logoalt Hacker News

A_D_E_P_T • today at 9:23 PM • 3 replies • view on HN

Trading punches in the benchmarks with Mimo v2.6 and 6.1-Sol (both very cheap!), and decidedly inferior to Opus 5.5. I'm afraid this looks unimpressive. Rather comical that they're delaying its launch "for safety reasons".


Replies

asdfasgasdgasdg • today at 9:33 PM

From the charts, it's similar in intelligence and cost per task to both Opus 5.5 (high) and 6-Astra (max). It would be better if it were more intelligent and less expensive, but I don't see a reason to expect it to have better performance than models released around the same time.

godbox • today at 9:37 PM

To play the Devil's advocate, Claude loves chugging tokens, while Gemini appears to be quite a bit more conservative and efficient. I believe AA's price per task breakdown reflects this.

thereitgoes456 • today at 9:32 PM

What are you talking about? It’s comparable to Opus 5.5 on “high” (54 vs 53; $1.82 vs $1.99), crushes every model except the most modern OAI/Ant ones, has way lower hallucination than every existing model and probably broader support for multimodal like existing Gemini models. This is so ludicrously off base.