I found the whole section super interesting... I'll copy it here:
I used a lot of OpenAI models to try and complete this port. In total I did over $400,000 in API priced tokens with GPT-5.6 Sol and GPT 6 Astra. They wrote over 1.3m lines of Rust over multiple months of /goal loops and never got past like 84% compat.
When I saw how little my Claude Code limits were burning, I figured it'd be fun to throw Opus 5.5 at this. It had a working v0 in 10 hours.
I assumed it kept using the code the Codex models wrote. I was wrong. Opus 5.5 started from scratch. It got further than Astra in 1/10th the time.
I let it keep going, and it definitely did. Total token spend was ~$24,047 of API spend over 2 weeks. I was using my Claude accounts, and it worked out to somewhere between 925% and 983% of my $200 plan weekly limits.
Expensive, for sure, but not that bad considering how much work has went into typescript-go.
Are we sure that he spend that much or he calculated the API pricing and used a subcription to implement this?
> I assumed it kept using the code the Codex models wrote. I was wrong. Opus 5.5 started from scratch. It got further than Astra in 1/10th the time.
Ouch, that is brutal and honestly, quite embarrassing but confirms what I have been seeing for a while. Personally, I find output from current OpenAI models still very hard to parse (though it has gotten better vs the pre-trains from both labs in mid/late 2025), thus hard to truly understand, verify and get comfortable maintaining vs current Anthropic models. I do occasionally see a higher ceiling in well scoped tasks with OpenAI models at the cost of (frequently) deviating from the original prompt in (sometimes) very destructive ways.
Could be that this hard-to-read output doesn't just go over my limited capacity/skills but with current models can become simply impossible to untangle beyond a certain size even when one has (essentially) infinite resources via multiple subs and different models.
Would also work with my suspicions for why OpenClaw (mainly build with Opus 4.5 and its post-trains) has been this hard to truly "fix", requiring highly paid Nvidia engineers, multiple months, (literally) infinite resources from OpenAI including access to internal models and yet still holds records for CVEs. Heck, another one was found just 7 days ago after what I'd argue was one of the most extensive hardening sessions any piece of software has ever undergone.
Makes my (multiple) decisions to start from scratch more than once on a major reworking of the existing tabbing interface in Firefox a bit less painful. Learned with each, found gaps in my knowledge, thanked the amazing docs the Firefox devs have been maintaining for decades and while starting from 0 was painful, getting back to MVP is easier than ever. When I hit a point were I was starting to struggle to truly parse additions a model was making to the patches applied to Firefox source code (even if they worked), I always found that pushing even slightly beyond that would incur painful, but hard to notice regressions, introduce major DB maintenance burdens as some models struggle to understand that in development regressions and incompatibility are acceptable and schema transitions aren't needed pre-release (still a case with GPT-6.1 Sol, less so post Fable for Anthropic), make me uncomfortable concerning privacy/security/data loss prevention (especially as I have seen Fable 5 cut some privacy/proper data removal corners in simple CRUD) and simply take away my control about the implementation. Wouldn't feel right to release something in that state, what I have now is fully understandable and thus could be maintained even without models.
Still expecting bugs of course, massively dreading security findings or even worse, possible data loss given browsers handle some of our most important personal+professional data and will surely have taken some embarrassing approaches that might have a much more performant solutions when implementing an infinite canvas of webpages, but still, rather that then also knowing I wouldn't even know where to start understanding a feature.
More so if, like with ts-rust, even (nearly) infinite tokens couldn't get me unstuck.
Great advertisement for Claude Code...