Reading the code does not mean you understand the code. One lesson that experience in software gave me: I never understood the code. You think it works a certain way, until you find out that it doesn't.
What LLMs make possible is for me to say: find out all the ways this thing works. Analyze the different ways we can run this software, build a fuzzer, build property tests, and run this software in every scenario possible. Log full traces. Log all the outputs. Now, analyze each scenario for bugs. You can't do that by hand.
If we are committed to it, if we put the resources towards it and dedicate the time to it (and we could do this just by saying: it will take half as long as it used to take!), software built by llms in healthcare, finance, automotive, defense, power plans, aviation, manufacturing can all be made MORE reliable and better with LLMs... without ever reading a single line of code. The LLMS are very good at logic, by the way.
Anyway all of this reads like someone who is not actually using LLMs to build software or hasn't tried them in a while. I felt the same way in 2025. I've written 100s of thousands of lines of difficult code. You, the person reading this, has probably interacted with software I've written. For a time you would've interacted with it every time you made a debit card transaction in the united states, for example. I understand code, and care about quality, and that's why I'm all in on LLMs for code.
> Reading the code does not mean you understand the code.
Reading the code may not be enough to understand the behaviour of your program, but believing you can understand the behaviour of a program without at least reading the high level code is truly silly.
(by high level, I mean the code living in the higher layers - of course we don't often read the code of the generated assembly, or the interpreter, or the browser, but that's because they're reliable abstractions, unlike prompts!)
Strange that the world worked before 2024 and software gets worse now. Your debit card transactions for example worked.
This sounds like a typical testimonial whose mind has become captive to Claude. It is like Scientology.
I agree, I think this false dichotomy between using LLMs and caring about quality/reliability needs to stop. All these mission critical industries listed in the article rely on extensive testing for quality assurance, with human code review being a layer on top of all that, but far from the most critical one.
Interpretability is the same, our abilities to do that have increased rather than decreased. I think a codebase generated by AI is actually more understandable than one generated by humans at this point, and you can ask clarifying questions whenever you get stuck.
TFA's points only make sense if the mental model the author has in mind is someone who writes a prompt then immediately puts an app into production without any thought behind it.
I don't think the discrepancy is in LLM capability improvements over the past year.
Correctness has never been a priority across an industry where rapid iteration and feature delivery drive sales. There's always some opportunity cost to doing things right, at the price of technical debt down the road. If AI is primarily used to produce fragile code, people will be wary of AI solutions. There's also ongoing public debate about AI safety and alignment. Deploying AI in safety critical applications feels riskier than ever in the current environment, even though it doesn't have to be.
For me, coding is like writing. The act of doing it is how you reason out the problem. There’s a lot of magical thinking you can get away with in your head that doesn’t get properly tested until you write it down. For me, vibe coding is great and fast, but I’m not getting the same opportunity to think through the problem I’m trying to solve.
> The LLMS are very good at logic, by the way.
It's wild to read this stuff and then also deal with the constant headaches of day to day hallucinations when interacting with Claude et al.
> build property tests
Even for a narrow use like this, you need to audit the output and have the skills to know that it did the right thing. I've seen it before where you give an LLM what seems like a clear interface and ask it write a test and it writes something shallow that doesn't actually test anything, or has serious problems.
> find out all the ways this thing works.
Mmm..aren't LLMs bad at exhaustively iterating all possibilities? So shouldn't the generated possibilities be manually checked?
You can provide the list of possibilities and use LLMs to generate the tests. Then you have to review the generated tests...
> I never understood the code. You think it works a certain way, until you find out that it doesn't.
Those are two separate claims, unless by the former you mean “I never perfectly understood the code.” You can understand code imperfectly. And even with LLMs, you can’t get truly infallible guarantees about a system.
> find out all the ways this thing works.
If you can’t understand the code, how do you know the LLM actually did what you asked it to do correctly? You wouldn’t know if it didn’t.
You somehow squeezed "you're holding it wrong" and "LLMs are actually good now" into the same sentence. I wasn't sure it could be done.
I think you can but you basically need to learn about a super simple and well characterised processor like the 8080 and write assembly for it. On x86/AMD64 there's no hope because they're out of order and have opaque instruction decoding. They could be doing anything! Performance and knowing what you're doing are sort of at odds with each other in that respect.
> I never understood the code.
> I understand code
Are you sure?
We could always build, more reliable software. The powers that be however (broadly generalizing) only care about solving the ‘problem’ superficially.
Therefore, we will end up with requests to do more (at the current level of quality/ reliability), as opposed to building better software
> What LLMs make possible is for me to say: find out all the ways this thing works. Analyze the different ways we can run this software, build a fuzzer, build property tests, and run this software in every scenario possible. Log full traces. Log all the outputs. Now, analyze each scenario for bugs. You can't do that by hand.
How do you know that it's verifying that the system under test exhibits the properties you desire without either understanding or making blind assumptions about the code it generates to build a fuzzer or a property test? It seems to me that you have just shifted the problem of verification elsewhere and introduced another potential source of error.
I’m using it a bunch. It saves a ton of time writing or reviewing code. It will catch things I won’t. But I’d express caution about the analysis or evaluation they do - LLMs will often confidently proclaim problems as solved or explain functionality and be wrong about it. Sometimes subtly, but sometimes just completely wrong. This is no different from humans, of course, except for the unabated confidence.
> I never understood the code.
> I understand code
Er, ok.
> Reading the code does not mean you understand the code.
Nor does writing it.
I hate that in our industry we refuse to operationalise reverse debuggers
"I never understood the code. You think it works a certain way, until you find out that it doesn't."
Bret Victor made a talk called "seeing spaces" in 2014 that should have woken up this whole industry: https://www.youtube.com/watch?v=klTjiXjqHrQ
He emphasizes that without the ability to see inside what is being built, creators often fall into "non-scientific thinking" (14:42), moving away from deep understanding and instead "blindly following recipes, from superstitions and rules of thumb" (14:47-14:51).
The worse is performance problems I've had engineers say some bizzaro things when discussing performance — we have the tools you can just measure the answer - we don't need to waste our time guessing
> I've written 100s of thousands of lines of difficult code.
How do you know it's difficult if you say you don't understand it?
> Anyway all of this reads like someone who is not actually using LLMs to build software or hasn't tried them in a while.
I run all the "latest and greatest" models the moment they become available to me. The amount of insanely bad code they produce remains largely the same, and largely in the same areas. And it cannot be caught by tests unless you know that bad code is there and end up with extremely bad tests anyway. I wrote about it here: https://dmitriid.com/adding-to-i-dont-read-ai-code-discourse
Main one is, of course, "to get a record from a database read all records from it, and filter in memory".
[dead]
[dead]
> I never understood the code. You think it works a certain way, until you find out that it doesn't. What LLMs make possible is for me to say: find out all the ways this thing works. Analyze the different ways we can run this software, build a fuzzer, build property tests, and run this software in every scenario possible. Log full traces. Log all the outputs. Now, analyze each scenario for bugs. You can't do that by hand.
Testing isn’t the same as understanding the code, or proving (even informally) that it is correct. Having the LLM do all these things above doesn’t lead you or the LLM to understand the code, to logically reason about its behavior over all possible states and inputs.
“Finding out that it doesn't” means that you didn’t properly reason through the code beforehand, checking all your assumptions against what the code and underlying systems are actually guaranteeing. This may be a matter of formal education (proving computer science theorems and algorithmic correctness in university), I don’t know.