logoalt Hacker News

spmurrayzzztoday at 3:34 PM4 repliesview on HN

> They do not beat opus on real-world usage

We have an internal eval that measures performance on tasks for a handful of embedded systems repos for our mmWave radios (mostly Rust, some C for microcontroller stuff). Qwen3.6-27B scores only 4% lower for pass@1, n=250 compared to Opus-4.8.

For the labeled dataset, the average PR size they're being measured against is around 1.5k SLOC.

This is very much "real-world usage" for us. The sort of change sets that come in daily/weekly and are solving non-trivial issues in the respective codebases.

As is usually the case, the most broad claims from both the labs and from the consequent pushback are talking past each other.


Replies

cyanydeeztoday at 4:36 PM

Let us know when you have Qwen vs Qwen comparison stats. As long as there's not a regression, that'd be awesome.

show 1 reply
HonshinMtoday at 6:02 PM

[dead]

enraged_cameltoday at 5:26 PM

>> We have an internal eval that measures performance on tasks for a handful of embedded systems repos for our mmWave radios

Okay but the parent said real-world usage, presumably meaning coding tasks.

We have a whole bunch of complex evals that Haiku 4.5 passes. That doesn't mean it is a good model for coding.

show 2 replies
croemertoday at 4:58 PM

How much does it score though? 0% would be 4% less if Opus was at 4%. Unless you mean relative fraction not percentage points - but people usually mean percentage points in such situations.

show 1 reply