logoalt Hacker News

epolanskitoday at 5:59 PM1 replyview on HN

+1, a single test means little.


Replies

jklmnopqrstuvwtoday at 6:05 PM

I don't think so. I specifically kept this PR to test model capabilities, and I've already tested a bunch of models. Current test results show that the more advanced the model is, the easier it passes. For example, GPT-5.5 Medium fails the test(has bug), but High passed.

show 2 replies