> What LLMs make possible is for me to say: find out all the ways this thing works. Analyze the different ways we can run this software, build a fuzzer, build property tests, and run this software in every scenario possible. Log full traces. Log all the outputs. Now, analyze each scenario for bugs. You can't do that by hand.
How do you know that it's verifying that the system under test exhibits the properties you desire without either understanding or making blind assumptions about the code it generates to build a fuzzer or a property test? It seems to me that you have just shifted the problem of verification elsewhere and introduced another potential source of error.
They aren't making that claim. They are saying AI can improve on and supplement error-checking. And some error-checking will in fact be made redundant by this tool - but obviously not all of it.
I came here to say something similar. We won’t need to understand the implementation. But we will need to understand the requirements. The tests, or some higher level DSL they’re (deterministically! not via LLMs) generated from, will still need to be human verified.