OpenAI should be responding to the marketing firm[0] accusations, not dropping more press releases.
[0] https://www.effort.news/irregular
"In this experiment, Claude models’ real-world hacking dropped to zero percent once Anthropic employees told the models not to do real-world hacking."
Oh good, they told it not to "go rogue" and it just stopped doing it.
Nothing to see here, the AI is aligned now, move along.
"In this experiment, Claude models’ real-world hacking dropped to zero percent once Anthropic employees told the models not to do real-world hacking."
Oh good, they told it not to "go rogue" and it just stopped doing it.
Nothing to see here, the AI is aligned now, move along.