logoalt Hacker News

manquertoday at 6:46 PM1 replyview on HN

The implicit point being adding this type of safeguards to Fable dumbs down the model in measured performance even though it is not fundamentally different.

Note it may not even be actual performance, typically in most benchmarks the model would be scored zero for refusing a task just the same as not completing it, so it could just be the Fable's stronger safeguards is just making it refuse more or perhaps even drop down to Opus.


Replies

Creamsicle47today at 7:34 PM

The model cannot complete that task, for one reason or another, and therefore it scores lower.