Hacker News
Stop Thinking of LLMs as Next-Token Predictors
LLMs are trained as next-token predictors during pre-training, adjusting probabilities toward tokens that actually follow given contexts in the data. Post-training methods such as reinforcement learning with verifiable rewards (RLVR) and RLHF let models generate new sequences, rewarding those that achieve desired outcomes, so the model learns from its own exploration rather than merely mimicking existing text. This shifts the role from pure prediction to goal-directed decision making.