logoalt Hacker News

schopra909today at 7:58 PM2 repliesview on HN

Progress is iterative. Everyone is always riffing on other’s ideas and can execute on them given enough support (eg $$). The person to get to an idea first is just 5% away, so it’s possible to catch up.

Moreover,I think it’s impossible to know if you’re hitting a portion of the sigmoid, because there will often be an idea that changes the trajectory altogether.

In 2024, there was a ton of talk about the plateau. Reasoning was an iteration on chain of thought, but it didn’t really work. Deepseek proposes RLVR as a way to get around the lack of $ they have to produce human reasoning trace data. That small iteration catches the eye of OpenAI and Anthropic, turns out to be way more important than even DeepSeek could have ever expected when it comes to improving LLMs for coding, and last 18 months have been an exercise on riding that insight to the nth degree.

That one small iteration brought us a lot of progress. Now we’re seemingly exhausting the impact of that one insight, but there may be another soon enough.


Replies

stymaartoday at 8:07 PM

> Deepseek proposes RLVR as a way to get around the lack of $ they have to produce human reasoning trace data.

What was the difference between what deepseek did for R1 and what OpenAI did for o1?

refulgentistoday at 8:54 PM

I don’t know why people think DeepSeek did reasoning models / RLVR before OpenAI, there was a gap of months.