logoalt Hacker News

When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

39 pointsby doppptoday at 4:10 PM52 commentsview on HN

Comments

otterdudetoday at 4:48 PM

This seems to be the end of the road for LLM's. There's only so much accuracy on a highly non-linear space you can get from a regression.

If the Pareto rule is any indicator, 80% of results come from 20% of causes. It seems that we have alot more to learn about intelligence.

I am reminded of a statement on truth from an ancient philosopher, this sentiment seems to be exactly the opposite of the LLM training paradigm

“The seeker after truth is not one who studies the writings of the ancients and, following his natural disposition, puts his trust in them, but rather the one who suspects his faith in them and questions what he gathers from them, the one who submits to argument and demonstration and not the sayings of human beings whose nature is fraught with all kinds of imperfection and deficiency. Thus the duty of the man who investigates the writings of scientists, if learning the truth is his goal, is to make himself an enemy of all that he reads, and, applying his mind to the core and margins of of its content, attack it from every side. he should also suspect himself as he performs his critical examination of it, so that he may avoid falling into either prejudice or leniency.” - ibn al-Haytham

show 11 replies
gertlabstoday at 5:13 PM

I started thinking about this back after the Llama 4 release, and since then our team has put a lot of thought into designing evaluations that don't saturate, are resistant to contamination, and can scale. What has worked best for us is using multi-agent environments with open-ended cooperative or competitive goals. Mostly designed as multiplayer games. The results tend to align with our experience for coding better than any non-aggregator benchmark, and likely at lower cost to run.

Data at https://gertlabs.com/rankings

show 2 replies
otterdudetoday at 6:11 PM

It appears y combinator has removed this interesting paper from its top trending position. Gee I wonder why?

show 1 reply
kanbankarentoday at 5:34 PM

37 authors and contributors need to be named up top?

Oh! I got my name on a paper! I don't think there is much reward for it these days.

hagen8today at 5:20 PM

Check out https://agents-last-exam.org/ there is still room for improvements!

tsunamifurytoday at 5:24 PM

I think its been pretty clear that in abnsense of clear use cases that are monetizable many model providers have been benchmaxxing on abstract or low utility average user performance.

This results in a lot of "oh wow it can do math I dont care about" and "it can't code a lot, but not well" outcomes instead of the core needs:

1) Cheaper faster and real time 2) Long walk capable without losing attention while rescoring goals over updated enviroment 3) Specific domain knowledge that can be trained quickly into the model (how we do work in this specific case)

behnamohtoday at 5:13 PM

This is AI slop. They didn't even change the plots default template.

buckle8017today at 4:57 PM

Slop

> We find that nearly half of the our bench- marks exhibit saturation

show 2 replies