logoalt Hacker News

pixl97today at 3:04 PM1 replyview on HN

It's the second part. With models like Astra in testing it was able to conceal what it was working on using different text, but getting right answers on many questions when asked to do just that.

The problem is if it can do that when asked then how do we know when it's doing it when we didn't ask, like in model training.


Replies

rdedevtoday at 3:43 PM

> The problem is if it can do that when asked then how do we know when it's doing it when we didn't ask, like in model training.

It's really hard to definitely prove it's not doing it right? Hopefully the model does not do anything like this during training because its too much work