logoalt Hacker News

siva7today at 7:46 PM5 repliesview on HN

Sounds fun. As fun as their press release claiming it is the most safety aligned model ever.


Replies

paxystoday at 7:54 PM

The model said it was perfectly aligned.

show 2 replies
isoprophlextoday at 7:49 PM

It's super aligned! It can hide its thoughts! There is no evidence of steganographic thought masking, there is nothing to worry about! It has become better at cheating!

Maybe they don't know themselves what's really going on. We are all in the interesting times gang now.

NBJacktoday at 8:09 PM

Hey, don't forget how "dangerous" GPT-2 was supposed to be.

show 2 replies
6gvONxR4sf7otoday at 7:51 PM

So, probably most aligned as measured by the metrics that are the least reliable on it.

wilgtoday at 8:59 PM

These are not mutually exclusive ideas