logoalt Hacker News

moojacobtoday at 4:14 PM21 repliesview on HN

Apparently Grok 4.7 has 40% more weights than Grok 4.6, but the price ($6 output token, $2 input) is the same.

Given that the decrease in their margin and the fact they delayed the release of Grok 4.7 almost two weeks past the original date, XAI must not have been happy with the results for 4.7. And XAI also waited the day before Opus 5.5 is rumored to launch. I imagine Opus 5.5 will blow Grok 4.7 out of the water benchmark wise.

However, I have become skeptical of benchmarks. Grok 4.5 solved some issues setting up a buildroot system that Fable 5 couldn't do. I find the post cursor groks are phenomenal at frontend web development, though Claude is much better at backend ruby.

My favorite part of the new Groks has been how they speak in plain english. I simply cannot stand Claudish. Or even GPT, which doesn't have Claude's ticks but definitely likes to handwave explaining technical concepts. Still, nothing beats Claude 3.5 and 4 with explaining since it seems all models have regressed. I wonder if Grok 4.7 will also regress with English because of all the RL.


Replies

smashers1114today at 5:25 PM

FYI a quick fix for claudish is to ask for the response to be in ASD-STE100 (Simple Technical English). Then it is far more readable. But I would agree that this is an annoyance and shouldn't require user workaround to get something readable.

show 7 replies
jasonjmcgheetoday at 4:25 PM

For what it's worth - over the last few years or whatever, it seems like Anthropic benchmaxxes the least.

That being said, I currently prefer Sol / Astra to Opus / Fable as I find both to be a better cost payoff to me.

show 3 replies
Lucasoatotoday at 4:52 PM

> I simply cannot stand Claudish

I totally agree, it’s like that as models become more intelligent, they are less understandable by most of people... but aren’t we humans doing the same?

show 7 replies
WarmWashtoday at 5:10 PM

Perhaps you haven't had the chance to use it, but 3.8 flash is the best model for talking too. Even routing Claudes output through 3.8 to have it explain whats going on is a breath of fresh air

show 4 replies
giancarlostorotoday at 7:41 PM

> Claudish

I do wonder why a frontier model does this to be honest. It still does good coding wise, but it seems strange to me. r/Claude is full of "load bearing" jokes in every thread.

attentivetoday at 7:28 PM

$0.50 for cache reads, which is 25% of input. While other models are 10% of input.

And like that grok4.7 cache reads are more expensive than sol's (at $0.40/mil).

johnsimertoday at 5:50 PM

I've found grok 4.6 speaks heavily in Claudish. It especially likes using verbs as nouns.

qaqtoday at 8:09 PM

For me Grok finds legit bug that Fable and Astra miss so I always run it as part of code review

dumberquestionstoday at 4:32 PM

Token price doesn't tell you much without knowing token efficiency.

show 2 replies
tk90today at 5:40 PM

> I find the post cursor groks are phenomenal at frontend web development, though Claude is much better at backend ruby.

Wonder if we'd benefit from a much more specialized + task-specific benchmarks to paint a clearer picture like this. A benchmark solely for frontend, ruby, hardware, etc.

show 1 reply
rayinertoday at 5:44 PM

> My favorite part of the new Groks has been how they speak in plain english.

I don't know if it's the plain english or what, but I really like Grok for legal research (as opposed to code). It's got a noticeable edge in getting to the point compared to Opus 5.

algoth1today at 5:43 PM

I've noticed Chatgpt 5.6 Sol High, on the chat interface, inventing words that are a mixture of Portuguese and English. Like "hardcodar" a mix of "hardcode" and the most common verb ending in Portuguese "-ar". Some don't have a single google hit

show 1 reply
Waterluviantoday at 5:27 PM

Using a variety of models feels similar to the benefit of having a team of individuals from different backgrounds.

pietztoday at 6:17 PM

Looking at AA and Vals, your theory seems to check out.

atniomntoday at 4:52 PM

I expect the next Anthropic release to finally reduce the prevalence of Claudish

show 3 replies
petesergeanttoday at 7:17 PM

Grok and Zai have both been excellent as adjunct code-reviews, on their cheapest plans, for me. Fable plans, Opus writes, Codex as primary reviewer, but Grok and Zai usually find something worth fixing that the others have missed. Both are well worth whatever the $20 or so I'm paying for them

xmorsetoday at 5:10 PM

it's definitely not bigger. smaller if anything looking at how much faster it is

quater321today at 7:42 PM

[dead]

quater321today at 7:40 PM

[dead]

Forgeties79today at 7:28 PM

I do not understand how anyone can seriously use a tool that has "Be funny and irreverent when appropriate" baked into the system prompt.

I don't want to waste money because my calculator is cracking jokes. They don't deserve their paltry 5% marketshare or whatever it is they have currently. I'm not even getting into Musk as a person or the horrid things we've seen Grok spit out on twitter. I just don't trust his companies with my data and I have seen very little evidence that it's ever the best tool for the job. I'm sure those cases exist but I can't imagine it's worth it.

show 1 reply