logoalt Hacker News

OsrsNeedsf2Ptoday at 12:26 AM0 repliesview on HN

I love obscure benchmarks, and I feel like I can trust their results a lot more - afterall, they (probably) weren't benchmaxxed. RuneBench[0] is another good example (how well LLMs can play Runescape)

[0] https://maxbittker.github.io/runebench/