Just read through some of the code "benchmarks", and I see why: https://reinvently.co.uk/tools/ed-o-meter/tests/
Most are extremely trivial tasks. I would be surprised if a model from 2 years ago failed these...
I think most of them would have actually failed, it's only recently that models were any good at using tool calls and harnesses after they started post-training for that
I think most of them would have actually failed, it's only recently that models were any good at using tool calls and harnesses after they started post-training for that