> Instead of finding a nerfed model, after six weeks of reconstructing wire logs, parsing transcripts, analyzing output tokenization, and staring at data, I found a much deeper issue. The model identity had remained the same, but the inference regime being delivered behind that model had not.
Is this something specific that shows up in the wire log, or is this the author's intepretation? The fact that Claude Code versions change over time in the test is suspicious. Anthropic has stated in the past that the underlying model behavior does not change over time, but Claude Code will change from version to version and this is expected. So if it's just Claude Code more aggressively tuning some knob in its requests, that's a pretty different thing than the underlying model changing.
They posted a long document explaining it all https://x.com/Lon/status/2101034933284417614
They're not measuring a fixed set of questions. This was post-hoc analysis on whatever prompts they were running each day.
Anyone can understand why it would go up or down depending on the work they're doing that day. This analysis is silly.