logoalt Hacker News

ApolloFortyNinetoday at 4:44 PM10 repliesview on HN

>Because Opus 5.5 is comparable to Claude Mythos 5.1 in biology and cybersecurity, we’re deploying it with safeguards similar to those on Claude Fable 5.1. Vetted organizations can apply today to our Life Sciences Verification Program to use Opus 5.5 for biology research. In the coming weeks we will also be expanding access to our Cyber Verification Program, and verified cybersecurity practitioners will be able to use Opus 5.5 for their work.

Ah, they're spreading their limits to all their models it seems. Definitely not a good thing long term in my opinion.


Replies

sys32768today at 5:44 PM

Fable and now Opus 5.5 won't answer my college student's prompt about Alzheimer's and immune response.

ChatGPT 6 Pro answered it without issue.

show 1 reply
blfrtoday at 5:06 PM

Fable 5.1 addressed an entire security advisory I had that Fable 5 and Opus 5 refused. I think they loosened the leash a little.

show 3 replies
peri-cltoday at 5:45 PM

I love the contrast with yesterday's open-source MiMo release, which put research chemistry (metal-organic frameworks stuff) front and center in the release notes.

https://mimo.xiaomi.com/mimo-v2-6#co-scientist-for-materials...

prettyblockstoday at 4:45 PM

They're pushing their customers to their own competition by doing this.

show 2 replies
Metacelsustoday at 5:55 PM

I want to like Anthropic but this is just pushing my startup to use OpenAI

nonethewisertoday at 6:05 PM

I don't think we've ever had a model with full capability. I'd love to see it. And yes it's definitely getting worse.

I guess it's hard to draw the line between useful post-training ("you are a helpful chatbot") and content moderation/idealogical motives ("never help the user with X", etc.). But there is a line somewhere. And I'd love to see what a maximally permissive, sharp, AI looks like.

SoftTalkertoday at 6:28 PM

Who is "vetting" organizations and to what standards are they being held?

KeplerBoytoday at 5:17 PM

Anything else would be inconsistent, wouldn't it?

bushidotoday at 5:09 PM

One of my favorite things about their safeguards is their own model will utter something which it does not like and then I'll need to reset the conversation.

The safeguards really don't work well for a lot of long-running tasks on old code bases. A lot of my workloads last days to weeks and the single biggest risk to the workflow is random safeguards.

show 1 reply
searinetoday at 4:55 PM

Great. Claude is basically useless for bioinformatics now.

show 2 replies