logoalt Hacker News

gpmtoday at 4:44 AM2 repliesview on HN

Yes?

If we look at the math problems they're solving their just now reaching the human frontier... they weren't doing that before.

And your comparison point is model released 2.5 months ago... saying for some use case you didn't see noticeable improvement in 2.5 months (even while other people and benchmarks disagree) isn't a great argument that they aren't improving.


Replies

contuberniotoday at 5:32 AM

Math problems are highly structured, very precisely defined, and already heavily studied and not very complicated compared to problems in engineering or finance. There's a lot of quality material on which to train and it's easy to tell quality apart from crap. The search spaces are a priori much smaller than in other areas and the people using the tools to study them are themselves good mathematicians.

Success in such problems does not automatically extrapolate to other contexts.

show 1 reply
jhrmnntoday at 5:17 AM

I think it’s more likely that that’s because no one tried to solve such problems with them before (OpenAI apparently started working in Navier-Stokes after a rumour that someone seriously advanced the problem with AI) plus improvements in orchestration. Fair, the latter could be as dangerous as stronger models.