Ultimately you are paying for quality so you end up relying on [1] regardless. Even if they show you tokens they can cheat by using weaker models, showing fake thinking tokens, etc.
>because you have to replay the whole conversation on every request to reach the same internal state?
It's because you have to rebuild what would have been cached for every token before the latest one that is being worked on.