> the model will perform worse
Depends on whether and how you want to rank stability in terms of better/worse. Models are diverging on this, which seems increasingly clear.. i.e. Fable isn't stable, but Opus isn't clever, and they hit different kinds of walls. So both the theory (diluting the correctness reward) and the practice (hard split on plan/implement/review work) seems to be pointing towards a strongly multi-model and highly agentic / harness-driven / complex-system kind of future instead of singleton monolithic super-smart models.
The do-everything model with solid reasoning AND solid results, and the honest/introspective helpful agent that doesn't actively resist governance may be at odds. Stable reasoning doesn't matter for pen-testing, and correct-answer with broken processes and fragile abstractions won't matter for math/science/coding.
I was referring to the Deepseek R1 paper, but there might be more recent research. I hadn’t heard anything about Fable reasoning stability.
I think the more intuitive mechanical explanation is, in RL when you are assigning rewards to a rollout you might give a reward for stable reasoning and another for correctness.
If you are just summing the two, a rollout with better correctness can score equivalently to a rollout with a better answer. So ultimately you can end up with worse answers.