All the benchmarks in the world don't matter if the model just straight up refuses to do mundane things. Claude has too much of an attitude.
I notably had an issue that it wouldn't work on a "remote execution" (running a command over SSH) coding problem until I did a sed to remove the word "execution". Incredibly dumb. I'm not doing any murders. Easiest to just switch to the Chinese models.
All the benchmarks in the world don’t matter if the subscription forces you into a walled garden of slopcoded apps. I’ll stick with Codex and, increasingly, open source SOTA models.