i guess to clarify my contention:
generic prompting like "review this code" or "make the architecture better" will raise the floor but cannot come close to human-quality code without humans understanding what code exists and what to ask for.
Obviously cursor and other labs have a stake in this going one direction, but I don't buy it, and I don't think you should either.
In fact, "Rebuild sqlite from spec" has all the problems with every other benchmark that I cited - model knows the whole problem up front and never has to iterate on the pile of slop it created cheating its way to a solution.
In any case, politely, I think we're mostly arguing vibes here and I'm not sure its going to get anywhere.
Fable writes code better than most humans can already so we've already surpassed human level coding. I think thinking otherwise is just coping for the job or industry.
And what percentage of the software in the world needs to come close to human-quality code? I’d argue the percentage of the whole is VERY low.