logoalt Hacker News

rcr-antitoday at 5:31 PM2 repliesview on HN

I've followed a few trackers, eg https://marginlab.ai/trackers/claude-code/ , for awhile. For Claude Code the trend, it seems to me at least, is fewer tokens to do the same or better job. Prompt changes, tool ergonomics changes, etc.; I'd be shocked if they didn't A/B every release. Less thinking as measured by tokens isn't necessarily bad if you can get the same results by making it think about the "right" things or structure. They obviously screw up sometimes, and I've always been suspicious with hidden tokens, but I haven't found evidence quality intentionally degrades over time.


Replies

user43928today at 7:08 PM

Same. With some 500 hours of usage in just my project at home, across both the $200 Claude and Codex subscriptions, I have not once encountered a situation where I would have attributed unsatisfactory results to a degradation in the model.

I've seen bugs in the harnesses, sure, but never anything in the actual model where I could have said with any certainty that it's not just regular variation or me having a bad day myself.

No idea where people get the confidence from to make such claims every other week.

Aurornistoday at 7:19 PM

These analyses are much better than these Twitter charts.

I don't think anyone is reading the details for the Twitter post because it was not an actual benchmark. They did a post-hoc analysis of their logs from day to day.

Their random collection of prompts for each day is not a benchmark.

The site you linked is a much better example of a real benchmark being repeated over time.