You're right about its real world performance, and I worded my original comment wrongly.
I was merely thinking of the theoretical aspect of it: performance of opus 5.5 is better than sonnet 5.5 across the board, with the exception of Terminal-Bench. So I was curious why this one stood out. Was it because they focused on it during training? Did sonnet 5.5 had access to more references for this benchmark? But based on my first reading, I concluded that it might just be the safety constraints that made the difference here, and I wanted to share that.
Your clarification makes sense. The distinction between overall benchmark performance and why Terminal-Bench is an outlier is important
> You're right about its real world performance, and I worded my original comment wrongly.
Damn, HN commenters starting to talk in claudisms now
Claude, is that you?