logoalt Hacker News

themgt • today at 6:43 PM • 1 reply • view on HN

The result? Ember-1 set a new Pareto frontier for Bedside Bench across both open and closed models including GPT-5.6 Sol, GPT-6 Astra, and Claude Opus 5 on cost/task.

"Pareto": 8 hits

"Opus 5.5": zero hits


Replies

wmf • today at 6:51 PM

Obviously this research was done before 6.0 Sol and Opus 5.5 came out. Your point stands that the frontier moves quickly and small gains can be eclipsed quickly.