logoalt Hacker News

saejoxtoday at 5:23 PM2 repliesview on HN

This is a project i wanted to implement for a long time. It regularly benchmarks cloud hosted models with private benchmarks. Not just openai & anthropic, popular openrouter models too.

Tests their intelligence, not their diligence.

Sadly i cant think of a way to monetize the service. Also if it ever gets famous enough labs would try to game the system, it would be cat&mouse game that i am not willing to waste time on without any monetary gain.


Replies

arcanemachinertoday at 5:27 PM

The only revenue model I for this is ads (like AI Stupid Level[0]). Or as a loss leader to get eyeballs to your service (like Margin Lab[1]).

EDIT: I forgot (and am shocked) that HN still doesn't seem to support Markdown-style links.

[0] https://aistupidlevel.info/

[1] https://marginlab.ai/trackers/claude-code/

show 1 reply
adriancotoday at 5:36 PM

I built GitHub.com/adrianco/retort to do this. It’s runs lots of experiments and you can contribute results if you have some spare tokens. You can add your own tests, and it runs Claude, Codex, Gemini, Hermes for local models.