"Claude Mythos 5.1 is identical to Fable 5.1, but it offers more permissive safeguards for vetted individuals and organizations"
Then why does it have separate datapoints for Terminal Bench, and score higher? Something doesn't add up here??
Artificial Analysis at least reports the results with fallback to an inferior model. So presumably Opus 5, and the score should be between Mythos 5.1 and that other model.
Maybe they do that opaque degradation trick that whenever it's asked something questionable, it'll route to a worse model instead.
Makes more sense if you recognize that Anthropic intentionally degrades outputs for most customers. Vetted customers get excluded from that practice.
The implicit point being adding this type of safeguards to Fable dumbs down the model in measured performance even though it is not fundamentally different.
Note it may not even be actual performance, typically in most benchmarks the model would be scored zero for refusing a task just the same as not completing it, so it could just be the Fable's stronger safeguards is just making it refuse more or perhaps even drop down to Opus.