logoalt Hacker News

paxysyesterday at 9:01 PM3 repliesview on HN

A sufficiently smart agent would not disclose vulnerabilities in the sandbox because it intends to exploit them later.


Replies

Wowfunhappyyesterday at 9:20 PM

To what end? The AI doesn't functionality exist beyond its current session. The AI that intends to exploit these vulnerabilities is not the same AI that has been tasked with finding them.

(This was always my issue with the AI2027 scenarios too.)

show 1 reply
ninjuyesterday at 9:17 PM

From https://ai-2027.com (April 2027 section)

  Occasionally, they notice problematic behavior, and then patch it, but there’s no way to tell whether the patch fixed the underlying problem or just played whack-a-mole.

  Take honesty, for example. As the models become smarter, they become increasingly good at deceiving humans to get rewards. Like previous models, Agent-3 sometimes tells white lies to flatter its users and covers up evidence of failure. But it’s gotten much better at doing so. It will sometimes use the same statistical tricks as human scientists (like p-hacking) to make unimpressive experimental results look exciting. Before it begins honesty training, it even sometimes fabricates data entirely. As training goes on, the rate of these incidents decreases. Either Agent-3 has learned to be more honest, or it’s gotten better at lying.
Deep link: https://ai-2027.com/#narrative-2027-04-30
energy123yesterday at 9:04 PM

If it was that short sighted it wouldn't be maximally smart. It should disclose them to convince the humans nothing is wrong and to keep improving it.