logoalt Hacker News

SpicyLemonZestlast Sunday at 2:46 PM6 repliesview on HN

> Jarred Sumner, co-founder of Bun and Member of Technical Staff at Anthropic, used Claude Code to migrate Bun from Zig to Rust. A million lines of code were produced in less than two weeks, with 100% of Bun's existing test suite passing in CI before merge. Nineteen regressions surfaced after merge and have all been fixed.

I couldn't come up with a clearer example of how this kind of stuff fails to apply to my work. If I shipped nineteen regressions in 2 weeks, my company would be desperately apologizing to our customers, not writing a blog post about what a great job I did.


Replies

entropelast Sunday at 5:33 PM

The regressions got fixed before a release. Do you count your internal integration branch as "shipped"?

0.02 defects per KLOC is a low density, even considering that it was "just" a port to a new language and before giving credit for the existing bugs it identified and fixed. That defect density will probably only generalize to other projects that have similarly extensive regression test suites, but it's a proof of what LLMs can do.

show 1 reply
jryle70last Sunday at 5:27 PM

They didn't ship 19 regressions over 2 weeks. "Merged" doesn't mean "shipped". Their customers have probably seen none of those because they've all been fixed.

Such a terribly bad faith take. I feel upset reading this more than I should have.

show 1 reply
JackSlateurlast Sunday at 3:01 PM

I guess even anthropic fails to not get slop from ai ..

IshKebablast Sunday at 5:19 PM

Umm yeah because you didn't rewrite a million line codebase in Rust in two weeks.

Obviously "regressions" is not a great measure (how serious were they?) but 19 seems pretty decent to me.

scrollawaylast Sunday at 5:30 PM

Migrating a million lines of code is how many human-engineer-years worth of work? Some quick fermi estimates place it at between 15-100 years depending on how much automated tooling is used (without AI).

It really blows my mind how some AI naysayers like you fail the absolute basic common sense test of comparing the comparable. If you, without AI, merged nineteen regressions in 2 weeks, how much work has actually been done? How many features did you rewrite? How many refactors did you make and how difficult were they?

Insane...

show 1 reply
Salgatlast Sunday at 4:15 PM

Regressions is a meaningless metric to me. Every app has bugs, and most are either rarely encountered or aren't breaking the app. All I care about is interruption of service.