The result? Ember-1 set a new Pareto frontier for Bedside Bench across both open and closed models including GPT-5.6 Sol, GPT-6 Astra, and Claude Opus 5 on cost/task.
Obviously this research was done before 6.0 Sol and Opus 5.5 came out. Your point stands that the frontier moves quickly and small gains can be eclipsed quickly.
Obviously this research was done before 6.0 Sol and Opus 5.5 came out. Your point stands that the frontier moves quickly and small gains can be eclipsed quickly.