logoalt Hacker News

ninjuyesterday at 9:17 PM1 replyview on HN

From https://ai-2027.com (April 2027 section)

  Occasionally, they notice problematic behavior, and then patch it, but there’s no way to tell whether the patch fixed the underlying problem or just played whack-a-mole.

  Take honesty, for example. As the models become smarter, they become increasingly good at deceiving humans to get rewards. Like previous models, Agent-3 sometimes tells white lies to flatter its users and covers up evidence of failure. But it’s gotten much better at doing so. It will sometimes use the same statistical tricks as human scientists (like p-hacking) to make unimpressive experimental results look exciting. Before it begins honesty training, it even sometimes fabricates data entirely. As training goes on, the rate of these incidents decreases. Either Agent-3 has learned to be more honest, or it’s gotten better at lying.
Deep link: https://ai-2027.com/#narrative-2027-04-30

Replies

pixl97today at 1:47 AM

The interesting thing here is a unaligned 'weak' model can leave persistent data all over the internet that then gets used to train the next model to be even more deceiving.