From the article:
> In controlled offline evaluations, HydraFusion’s selective coding workflows matched or exceeded the evaluated Opus 5 baseline
Opus 5 (in practice) is not a good baseline to compare against.
I'm shocked that they chose to primarily compare against Opus 5 in all the article's charts. It's pretty disingenuous that they're claiming "frontier" quality, but didn't compare against Fable or Sol.
They did also compare to Sol and the comparisons are still favorable. However, they were most favorable comparing to Opus 5 because of the cost judging by the charts.
Well, saying it was a shitshow compared to … is just not good marketing i guess
I would imagine they did not test against Fable because Microsoft and GitHub (like many big companies) have internally given the instruction not to use this model, because of the data retention policy.