logoalt Hacker News

mattlondontoday at 3:42 PM13 repliesview on HN

Currently top at https://deepswe.datacurve.ai - beating Opus 5!

https://artificialanalysis.ai/models/gemini-3-8-flash shows an intelligence score of 59, the same as Opus 5 medium!

Wow - for a flash model this seems to benchmark powerfully. Remains to be seen what it is like to use.


Replies

theHocineSaadtoday at 4:08 PM

As of writing this comment, Claude Opus 5 has an intelligence score of 63, not 59 (it's not the same as Gemini 3.8 Flash).

With a score of 59, Gemini 3.8 Flash is in eighth place, falling behind even Grok 4.6, Kimi k3, and GLM 5.3.

https://imgur.com/a/BMOJBED

show 3 replies
markasoftwaretoday at 3:54 PM

On artificial analysis it's only equal to opus 5 medium effort. Opus 5 max scores 63.

Further, opus 5 medium outputs 4x fewer tokens to achieve the same result, negating a lot of the speed difference.

show 1 reply
WarmWashtoday at 3:47 PM

The benchmark also doesn't include speed. You almost think something has gone wrong when using it because it returns full responses so incredibly fast.

show 1 reply
onlyrealcuzzotoday at 3:53 PM

The rumor is that 3.9 is an equal improvement in all directions, and that it should be another fast follow on like 3.7 and 3.8 were.

It's almost across the board better than Terra at less than half the price. 3.9 is likely to approach Sol at the 1/10th the price.

Hopefully OpenAI releases Astra first, and it's not only better than Sol but significantly cheaper, too.

show 1 reply
bertilitoday at 4:09 PM

A fifth of the cost of Opus 5! Google is certainly pushing the completion with this.

show 1 reply
ttultoday at 3:49 PM

Crushing it on DeepSWE is a very big deal. Excited to give this a try.

show 2 replies
Gecko4072today at 3:43 PM

Google - we're so back

show 1 reply
kimjune01today at 4:36 PM

deepswe is public and can be considered contaminated.

sunaookamitoday at 3:50 PM

>shows an intelligence score of 59, the same as Opus 5!

...on Medium reasoning. Claude Opus 5 (high) is the default in e.g. Claude Code and scores 61. Still very impressive.

WhitneyLandtoday at 4:41 PM

There are important gaps in that hot take.

For example, it's not even close to Opus 5 on Terminal-bench 4.0, 19.1% vs. 51.8%.

satvikpendemtoday at 3:50 PM

We'll see about that. I suspect benchmaxxing as all the labs do as I haven't found Gemini models to be nearly as good in agentic engineering compared to Claude or GPT models.

show 2 replies
jrflotoday at 4:24 PM

sidenote, but wow sonnet 5 is shockingly bad on this benchmark.

show 1 reply
pkos98today at 4:11 PM

Wait a week with your judgement - most likely, Google is just bench-maxing very hard. If you look at the previous Flash models and the announcement on Google I/O, it was an absolute disaster. Reality diverged very much from the marketing (supposedly great benchmarks).