logoalt Hacker News

whinvikyesterday at 6:10 PM2 repliesview on HN

Yeah either the benchmark isn't very useful anymore or V4 Flash is a really, really good model.


Replies

fallingbanannayesterday at 7:37 PM

GPT 5.6 Luna is an extremely cheap and still very capable model.

A chinese model being in the same ballpark of capability at half the price sounds believable to me.

ignoramousyesterday at 6:28 PM

In my use, DeepSeek v4 Flash (which replaced the quite excellent MiniMax M3) lags behind GLM 5.2 & Muse Spark 1.2 (let alone Kimi K3). Also, K3 is a much bigger multi-modal model, while Flash is text-only and likely optimised for coding tasks.

show 1 reply