Because RLHF causes the opposite effect. RLHF is how we got the wave of "AI Psychosis" in 2024-2025, because the models never disagreed with people.
That whole episode caused the whole industry to shift away from RLHF, and towards RLAIF, RLVR, and DPO, and add a lot more safeguards, tests, and reward functions that push models in the direction of doing the opposite of what people want and confronting and strongly correcting their users, if it has determined the user is wrong.
Because RLHF causes the opposite effect. RLHF is how we got the wave of "AI Psychosis" in 2024-2025, because the models never disagreed with people.
That whole episode caused the whole industry to shift away from RLHF, and towards RLAIF, RLVR, and DPO, and add a lot more safeguards, tests, and reward functions that push models in the direction of doing the opposite of what people want and confronting and strongly correcting their users, if it has determined the user is wrong.