logoalt Hacker News

DetroitThrowtoday at 5:15 PM1 replyview on HN

I've tried it on my "let's run every model in parallel and see which finds more edge cases" type of tasks, and Grok 4.5 was really behind Opus/ChatGPT but ahead of Gemini - despite having a strong showing on benchmarks.

That makes me really skeptical of it being GPT5.6-tier, much less Fable-tier, based on some of these benchmarks alone. But I'll test here shortly.


Replies

DetroitThrowtoday at 6:24 PM

It's still not as good as GPT5.6 or Opus5 but it's better than KimiK3. Good job xAI team.