Pretty crazy that the model doesn't know that it needs to stop before it hits 128k output tokens. I guess it has no sense of how many tokens in it is? Wouldn't this be possible to work into the architecture?
I think this is a bug. I've not seen this problem from any of the other frontier models.
I think this is a bug. I've not seen this problem from any of the other frontier models.