Pretraining
How does it learn?
The longest and most expensive stage: the model reads enormous amounts of text and adjusts its parameters to get better and better at predicting the next token. That is where language, facts and styles come from.
An image to remember it
Someone who read the whole library but never had a conversation. Knows a great deal and has no idea how to help anyone.
The misunderstanding
It was programmed with rules. Nobody wrote rules like "if they ask X, answer Y". There are patterns learned from the data.
What changes in practice
That is why it sounds confident about any topic: it learned to produce plausible text, not to check it.