Stage 2 · What happens when you write to it
Next-token prediction
How does it generate an answer?
Generating means repeating one step: look at the whole context, pick a likely token, add it and start again, until the answer is finished.
An image to remember it
Your phone's autocomplete, with far more context and far more training.
The misunderstanding
First it thinks up the answer, then it writes it down. It writes it token by token. That's why asking it to reason before concluding, or using models that "think" first, improves the results.
What changes in practice
The same question can get different answers. That's normal; if you need consistency, set the format and the criteria in the request.