What about multi-token prediction and speculative diffusion? That’s a different mechanism of prediction, even if it serves only to accelerate decoding.
As you say, that's just an efficiency play and, as I understand it, doesn't change the behavior of the models beyond perhaps a small amount of sampling noise.
As you say, that's just an efficiency play and, as I understand it, doesn't change the behavior of the models beyond perhaps a small amount of sampling noise.