This doesn't look like a plateau to me: https://artificialanalysis.ai/evaluations/artificial-analysi...
I do agree that they're investing heavily in brute force methods though. I've been trying out GPT-5.6 Sol "Ultra" recently and that thing fires up a bunch of subagents and crunches for hours.