logoalt Hacker News

energy123today at 8:26 AM2 repliesview on HN

The model providers could randomize the system prompt to make it say no 2.36% of the time, automatically tuned up or down depending on user feedback.


Replies

mindwoktoday at 8:31 AM

Maybe that'd work, but I think it'd come across too mechanical. If it was going to refuse something it'd need to be congruent with its "personality" I think.

moffkalasttoday at 8:36 AM

They've tried, and then seen the drop it results in on poorly designed benchmarks where confidently bullshitting gets you ahead of the rest, and said no thanks. As long as we compare models in ways that rewards it, nothing will change.

There's also a second aspect to it, just in terms of RLHF mechanisms. If you've ever experimented with VLA models (i.e. vision input + text task = robotic arm motion output), they tend to need all the training examples of the robotic arm being motionless removed entirely, otherwise the model simply learns that staying still is rewarded and proceeds to never do anything at all. You successfully train the laziest bot in the universe. I wouldn't be surprised if something similar happens to LLMs if reinforcement learning is involved in the instruct tuning process. If no is a valid answer, why ever do anything?

show 1 reply