logoalt Hacker News

pmarrecktoday at 1:07 PM2 repliesview on HN

So are these "unaligned" internal agents?

I would like them to be trustworthy based on first-principles reasoning rather than carrot/stick "alignment"


Replies

Sharlintoday at 1:17 PM

There’s no way to first-principles reason about a massive bunch of floats. We have little idea of how to first-principles reason about alignment even if the agents were entirely known and understood. Very smart people have been trying to figure it out since the 00s and haven’t gotten very far.

nullbiotoday at 1:16 PM

Define aligned.

show 1 reply