The most important question is how much more unreadable the code became after all this "ratcheting the benchmark down". If you unroll a loop it will perform faster, but making changes to such unrolled code will be a mess. Will this make them ship slower overall? I'm sure at least half of it was just poorly written React code, but the other half?
It's the same problem as overfitting in model training. If you're not measuring something it will get sacrificed.
Or, perhaps the code quality literally doesn't matter anymore and we've reached "code quality escape velocity" where you can code as much slop as you want, the next generation of models will clean it up faster than the slop generates?
Came here to post this. AI rationalists love thinking about paperclip maximizers, but don't seem to care when it turns their own codebase into paperclips. Or to rephrase, turns their whole engineering org into meat proxies, slowing down engineering productivity in the long term because understanding is drained out of the staff and flushed down the drain every time they close a Claude Code tab.