I doubt that would change the perception. Every model release is followed by accusations of nerfing.
There are several projects that repeat benchmarks on published models. None has ever found significant fluctations
Here's one example https://marginlab.ai/trackers/claude-code/
Fluctuations of a few percentage points are to be expected and should not surprise anyone who knows how LLMs work.
This Twitter analysis of Fable 5 is not that at all. They analyzed their coding sessions and blamed all of the fluctuations on Fable changing. They then compared to ARC-AGI-2 questions as the benchmark for thinking tokens and tried to stir up anger that coding turns don't produce as many thinking tokens as the ARC-AGI-2 problems.
This page has been in 'New model — collecting baseline data. Degradation detection paused.' state for months now. It seems to never say 'degraded'.
If you look at the graphs, the latest benchmarks are showing a pretty significant dip, and they match pretty well with some horrible experiences I've had in recent weeks. You can see token usage steadily going down, matching exactly what the author measured on his own.