logoalt Hacker News

fragmedetoday at 7:13 PM3 repliesview on HN

We could still have soft evidence though. Make a Todo app on Monday, and make a Todo app on Tuesday, and see what it makes in comparison.


Replies

ArvidSutoday at 7:45 PM

You only need to come up with a catchy "SomethingBench" name, post it on reddit/x and now you're an ai sage. Not to disparage the launch/after comparison though, I'd genuinely enjoy a data point like that

marcus_cemestoday at 7:41 PM

You would need a significant sample size to make any sort of conclusion from such a probabilistic process. Then there's the issue of how you would actually grade/compare.

luckydatatoday at 7:50 PM

someone already does that https://aistupidlevel.info/

show 1 reply