Sounds fun. As fun as their press release claiming it is the most safety aligned model ever.
It's super aligned! It can hide its thoughts! There is no evidence of steganographic thought masking, there is nothing to worry about! It has become better at cheating!
Maybe they don't know themselves what's really going on. We are all in the interesting times gang now.
Hey, don't forget how "dangerous" GPT-2 was supposed to be.
So, probably most aligned as measured by the metrics that are the least reliable on it.
These are not mutually exclusive ideas
The model said it was perfectly aligned.