logoalt Hacker News

WithinReasontoday at 10:52 AM6 repliesview on HN

Mixed signals, here it's performing below even GPT-5.4 Nano:

https://livebench.ai/

while here it outperforms Fable by a significant margin:

https://oxalpha.com/

but if the latter is true, will people still say it was "distilled" from Fable?


Replies

woadwarrior01today at 11:27 AM

That benchmark is super sus. Until someone pointed it out, the top performing open weights model was a Kimi K3 fine tune from their sponsor (abacusai/Smaug-Agentic). Now, it's not on the list.

Source: https://twitterwebviewer.com/?tweet=2091116504787935350

Aurornistoday at 12:16 PM

Claims about Ox Alpha performing at Fable level were from the social media hype cycle. Everything new in the LLM space brings a wave of influencers hyping it up as a revolutionary leap forward. Don’t forget to like and subscribe to learn more.

It is a capable small model, but it’s not frontier level. The interesting part will be seeing the model size, how it responds to quantization, and how fast it runs on the kind of non-server hardware that we can buy without selling a kidney.

show 1 reply
sunbumtoday at 10:56 AM

the 2nd website is not official, just something someone slopped together for some reason.

show 2 replies
tescrealtoday at 11:34 AM

I really want to see hard evidence of distillation before I buy into it. Seems like a lot of sour grapes over not having the sort of lead assumed. In this field, it has been shown repeatedly that leaps in performance come swiftly and without notice.

show 2 replies
epolanskitoday at 11:02 AM

GLM 5.3 was a great model, so this would be strange to release a regressed model

show 1 reply
re-thctoday at 11:13 AM

the outperform Fable was a mid (not completed) benchmark run. Real results were lower.