logoalt Hacker News

kadobantoday at 2:33 AM1 replyview on HN

Oh, wow, they think it's just a smidge below the q4? That's crazy good if true.


Replies

anana_today at 3:03 AM

The benchmarks they chose are rather cherry picked to not include long context or difficult ones that involve long horizon work or many agent turns, as I suspect this is where the model shows more differences compared to the full fat one

show 1 reply