According to benchmarks in the announcement, healthily ahead of Claude 4.6. I guess they didn't...

Metacelsus • yesterday at 5:05 PM • 7 replies • view on HN

According to benchmarks in the announcement, healthily ahead of Claude 4.6. I guess they didn't test ChatGPT 5.3 though.

Google has definitely been pulling ahead in AI over the last few months. I've been using Gemini and finding it's better than the other models (especially for biology where it doesn't refuse to answer harmless questions).

Replies

CuriouslyC • yesterday at 6:38 PM

Google is way ahead in visual AI and world modelling. They're lagging hard in agentic AI and autonomous behavior.

throwup238 • yesterday at 5:22 PM

The general purpose ChatGpt 5.3 hasn’t been released yet, just 5.3-codex.

neilellis • yesterday at 5:45 PM

It's ahead in raw power but not in function. Like it's got the worlds fast engine but one gear! Trouble is some benchmarks only measure horse power.

➕ show 1 reply

scarmig • yesterday at 8:29 PM

> especially for biology where it doesn't refuse to answer harmless questions

Usually, when you decrease false positive rates, you increase false negative rates.

Maybe this doesn't matter for models at their current capabilities, but if you believe that AGI is imminent, a bit of conservatism seems responsible.

Davidzheng • yesterday at 5:54 PM

I gather that 4.6 strengths are in long context agentic workflows? At least over Gemini 3 pro preview, opus 4.6 seems to have a lot of advantages

➕ show 1 reply

nkzd • yesterday at 6:41 PM

Google models and CLI harness feels behind in agentic coding compared OpenAI and Antrophic

simianwords • yesterday at 5:06 PM

The comparison should be with GPT 5.2 pro which has been used successfully to solve open math problems.

alt Hacker News

Replies