Nope, it doesn't.
No logic required, you can just build an LLM yourself, including post training. You'll see that predicting the next token isn't something the model does or is optimized for in RLHF or RLVR. You can hand wave all you like, but you have never done it.
Yes, no logic is necessary for LLM adherents we're all finding out.
Carry on good soldier.