I don't think I learned a single thing from that entire article.
I might be in a bad mood or just am already familiar with the subject (from reading hacker news frontpage??) but the gist of the article is that we stopped doing things with one agent and now we do things with multiple agents and this is somehow (unspecified) better.
For problems that are limited by tokens/second, well, this is rather self obvious how it helps, for other problems... what? He made a big deal about agents "sending messages" but how is that functionally different from the internal transformer/attention operations except now operating over tcp?
This rhetoric (and it is rhetoric, in that it is designed to persuade) too often lacks nuance. Generating a 3 minute stand alone video is impressive, but it is a different task from generating code that must be consistent with an existing mountain of code. Generating code that implements a straightforward feature within an existing system is different to generating code that must make architectural decisions in novel domain or with a novel architecture. The Navier-Stokes result is impressive but 1) they spent a bajillion tokens / currency units to achieve it and 2) AFAIK all the recent mathematics results have been counterexamples, which are a different beast to proofs.
My own experience, pootling around with Claude like a peasant, is that it is simultaneously amazing at some tasks and an absolute moron at others. The more abstract the reasoning required, the worse it is.