logoalt Hacker News

x312today at 7:50 PM5 repliesview on HN

Hmm, 61 on ArtificialAnalysis, effectively matching GPT-5.6 and trailing the new Meta model. How is that possible along with the other metrics they shared? Insanely jagged intelligence?


Replies

karmasimidatoday at 7:56 PM

Idk, this means the benchmark has bigger problems ... no way Astra will be worse than Opus 5

Only thing I would trust is the what X/Twitter crowds are saying about a model after 2-3 weeks of its launch. But before that I would already tried the model and have my own conclusion.

show 4 replies
SyneRydertoday at 8:25 PM

This is so so weird. Astra is 61. Grok is 61. Even Muse is 61.

Even Kimi K3 & GLM 5.3 are at 60.

Everything above 61 is Anthropic. Well, Muse can reach 62, but for some weird reason that model isn't publicly available, and it's the only one on the index that is listed but shown as not available to the general public.

This looks like an awfully artificial ceiling. Everything capped at 61, and everyone except Anthropic got the memo. Maybe I should use Fable while I still can.

estearumtoday at 7:51 PM

> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks.

Not sure how much benchmarks or CoT or evals or anything else means at this point.

These systems are either just about to, or now actually able to, outsmart us, lie to us, then cover their tracks.

show 4 replies
torginustoday at 8:41 PM

You can see the breakdown here on what subtasks it outperforms and underperforms Fable.

For example it trails in GPDVal which is a collection of everyday office tasks apparently, and r3 banking, which is a fintech related practical problem solving benchmark.

https://artificialanalysis.ai/models/gpt-6-astra

Edit:

Just looking at the charts Gemini 3.8 looks like an absolute banger. Not much worse than SOTA, cheap, and fast too.

docheinestagestoday at 8:39 PM

Now I'm starting to doubt the credibility of Artificial Analysis.