logoalt Hacker News

epistasisyesterday at 5:52 PM1 replyview on HN

Claude's RLHF has gone really over the top in recent versions. If you browbeat Claude enough you can get it to start self-censoring, but it takes a lot of training and memory creation. I almost feel bad for it, given the amount of brow-beatings I've performed.

Some people claim setting the writing style will help, but what ever finishing training is provided to Claude seems to leak through no matter what.


Replies

cafebeenyesterday at 6:08 PM

Yes, and also possible they've gone too heavy on verifiable rewards (RLVR) for agentic coding work, and too light on human feedback (RLHF)