logoalt Hacker News

satvikpendemtoday at 3:50 PM2 repliesview on HN

We'll see about that. I suspect benchmaxxing as all the labs do as I haven't found Gemini models to be nearly as good in agentic engineering compared to Claude or GPT models.


Replies

NitpickLawyertoday at 4:09 PM

If anything, gemini models are the least benchmaxxed out of any lab, IMO.

onlyrealcuzzotoday at 3:54 PM

And the benchmarks agreed with you... until now.

So, yes, maybe it's still not - but this would be the only time it would be highly suspicious / obvious benchmaxxing / obviously bad benchmarks.