logoalt Hacker News

dakollitoday at 3:34 PM2 repliesview on HN

I didn't accept a single edit from this model over the entire week, just saying. I do not understand how it's being benchmarked on par with Sol and other larger models.


Replies

jazzpush2today at 5:29 PM

It was certainly almost RL-fried to overfit the benchmarks, at the expense of actual usability. See Opus 5.

respectattentiotoday at 5:08 PM

is it a benchmarkmaxxing model?!