During training, certain tokens are more likely to lead to a lower loss function value, which is how you "win" the game of LLM output.
So, next-token predictors
So, next-token predictors