Because the incentive has been changed from the true best output always, to a mix of "close to the best but not always" output.
For the (majority) of us using Claude models for computing as a tool, obviously we're not going to be thrilled that our new tool will perform worse going forward.
> true best output always
literally never how it has worked
Do you understand that LLMs are probabilistic?
Ask a model the same question twice and you will get different results. So, how were you ever getting “the best result, always”?
Put the watermarked version head to head with the non-watermarked version.
If you can't tell which one is better then how can you make any assumption about performance?
For all you know performance is the same.
So many people complaining about something they quite literally have zero evidence for.