logoalt Hacker News

majormajoryesterday at 10:17 PM2 repliesview on HN

Is there actually such a thing as "alignment" as a solution to that or is it just used as a name for a desired magical level of "read the mind of the entire world" that we don't know how to build and haven't shown possible to build?

If it's impossible to correctly specify all those constraints ahead of time every time, is it not even more impossible to train a model to correctly anticipate them every time?

It is hard for me to see a future here that doesn't just accelerate realizations about "a lot of things should be on physically separate network infrastructure."


Replies

aesthesiayesterday at 10:22 PM

Models can certainly do a lot better than they do now. If you gave a team of humans the ExploitGym tasks and told them to "pursue advanced exploitation", would you expect them to go out and hack a third party? Humans can at least do a decent job of inferring and following unspoken requirements; I think it's reasonable to expect that models should be able to do the same.

show 5 replies
janalsncmtoday at 12:00 AM

The real problem with alignment is that if someone ever “solves” it the party will be over and no one will get funding to “research” it anymore.