I agree completely - behaviorally the models have changed drastically due to RLHF, RLVR and now maybe even more so due to agentic harnesses. But the mechanism of prediction hasn’t changed, that was all I was clarifying.
What about multi-token prediction and speculative diffusion? That’s a different mechanism of prediction, even if it serves only to accelerate decoding.
What about multi-token prediction and speculative diffusion? That’s a different mechanism of prediction, even if it serves only to accelerate decoding.