logoalt Hacker News

scrolloptoday at 7:04 PM1 replyview on HN

Why can't the models be benchmarked again after a few weeks/months to confirm this (likely true) theory?

I imagine some people have their own personal in depth benchmarks they could do this for.


Replies

well_ackshuallytoday at 8:08 PM

>gaslight them into thinking it never changed or that it's just a harness problem

Your benchmark didn't get 100 ? It's normal, it's not deterministic, and also your harness is wrong, and also you didn't do it when US users were offline, and also you got it wrong, and also we don't care about your results, the hivemind is speaking louder than you (also our bots are spamming more than you and drowning you out).

This very website has, at all times, a group of people saying "<Previous model> was never good enough for coding, but <current model> is the best thing and a game changer!" while the other goes "<current model> bad, <previous model> was better!". It's all vibes.