logoalt Hacker News

janalsncm • yesterday at 11:20 PM • 0 replies • view on HN

I was referring to the Deepseek R1 paper, but there might be more recent research. I hadn’t heard anything about Fable reasoning stability.

I think the more intuitive mechanical explanation is, in RL when you are assigning rewards to a rollout you might give a reward for stable reasoning and another for correctness.

If you are just summing the two, a rollout with better correctness can score equivalently to a rollout with a better answer. So ultimately you can end up with worse answers.