These tests are pointless when often models get nerfed few days after release. Astra today is way dumber than just few days ago.
Is that something we have credible evidence for? Do they serve a better model when artificialanalysis (the benchmark site) is making the requests, and so on?
Not Astra but Sol hasn't been nerfed: https://marginlab.ai/trackers/codex/
How and why do they get nerfed? To save money?
Is that something we have credible evidence for? Do they serve a better model when artificialanalysis (the benchmark site) is making the requests, and so on?