logoalt Hacker News

timmgtoday at 4:41 PM1 replyview on HN

Well: do they listen to the evaluators?

I think there are any number of reasons a "safety" person could be concerned about any state of the art models. And so I would expect at least one of those to apply to any model.

The question then is: do we stop when the safety people say to (they will) or not?


Replies

stratos123today at 5:34 PM

Yeah, I'm also pretty skeptical about this. With AI companies we see time and time again that they can have benevolent, well thought-out regulations and then a few years just... abandon them - the most notable case of this being, of course, the founding of OpenAI as a nonprofit dedicated to benefitting all of humanity, and it being stolen by Sam Altman.

If Anthropic just unilaterally does the evaluator thing and can't achieve cooperation of the rest of the plan, my guess is that it'll have some impact for a few months and then they'll just stop reacting to the evaluators' reports and the evaluators would stop bothering to report anything. I think the idea is that if Anthropic does get government support for this, the external evaluations will be legally binding. The problem with this, though, is that the current US government perhaps can't be trusted to consistently enforce a regulation on a company, rather than e.g. taking bribes to not do so.