logoalt Hacker News

How well do agents use test/verification techniques?

63 pointsby vinhnxtoday at 2:58 AM14 commentsview on HN

Comments

kqrtoday at 7:54 AM

I don't know what I'm most impressed by: (a) the testing expertise, (b) the effort spent looking into the reasoning mistakes generated by LLMs, or (c) the insane amounts of money this must have cost!

ivanzhaowy123today at 7:46 AM

In my experience, agents often think of more edge cases than humans when writing unit tests. But under the guidance of certain skills, they can become mechanical and lose sight of the business logic.

For example, when I use the Superpowers skill set, the agent proactively adopts TDD for every new feature. But its understanding of testing often stays superficial: if the user asks for a screen with a “Send” button, it first writes a test checking whether the button exists. The test fails, so it adds the button to make it pass.

As a result, the test suite fills up with low-value cases that check whether a property exists or a string matches exactly. The agent follows the “write a failing test, then implement the feature” workflow, but never really tests the business behavior: When should sending be allowed? What should happen after success or failure? How should duplicate submissions be handled?

The problem isn’t that agents can’t write tests. It’s that they seem prone to reducing TDD to a rigid sequence of steps, struggling to independently derive meaningful test cases from business requirements and use them to drive development.

quietrastertoday at 7:52 AM

nice setup for this. did any single technique clearly stand out, or did it mostly come down to the base model?

ngruhntoday at 6:42 AM

I wonder how well this fares

   try to make illegal states unrepresentable
In my experience, agents know how to do it. They just don't if it's not the default style of the language.
sisciatoday at 5:03 AM

It is still early, but I find that this experiment makes little to no sense and it is barely useful.

The way you test code cannot (and should not) be decoupled by the way in which you architect the code itself.

80%+ of effective testing is not in the testing framework but in the code architecture.

The author doesn't mention how the code is being architected and managed.

For what it is worth, I found that forcing agents on DI/hexagonal architecture and forcing a trivial coverage check is quite useful and produce overall good enough code with relative little effort

gz09today at 5:08 AM

Results seem somewhat reasonable given that the amount of verus/TLA/Creusot/Lean code out there is tiny compared to all the other non-formal code.

So it's understandable that the agents wont be able to go beyond proving trival things, given how much more difficult it is to write such code.

A (more) interesting experiment (to me) would be to write a high level spec manually for a non-trivial system (liveness etc.) and see if the agent can produce an implementation using guided refinements that satisfies this specification.

show 1 reply
tomrodtoday at 4:46 AM

Timely. I'm also looking at this now, actually! We tend to throw benchmark after benchmark at systems, but miss that models are one part of the system. Harnesses are more than models and need tuning too, and in doing so there can be gains or loss of prior tested function as well.

anitiltoday at 5:19 AM

You know how everyone thinks agents are bad at the thing they're good at and good at the thing they're bad at? It turns out I must be bad at testing, because I thought they do a reasonable job.

show 1 reply
vetronautatoday at 4:51 AM

If I understood correctly, the author had the agent perform manual mutation testing (write code, write passing tests, manually change the code to see some tests fail, then revert the change), rather than automated mutation testing. Why not use a mutation-testing framework and consider the build failed if a certain percentage of mutations are not killed?

The article claims that the agents didn’t actually use TDD or mutation testing. While it’s possible for an agent to ignore the TDD procedure even when instructed to follow it, it can’t ignore a build failure.

ianjbutlertoday at 5:16 AM

Detail in TFA looks impressive and I promise I'll do a close reading later. But..

The whole premise of the question is hilarious. They change text in an existing one-line comment and the best models in the world think, gee, maybe I'll lint everything AND run 4000 units. Maybe 5000 integration tests too, just to ensure we collide with any other work in progress. So you write the obligatory but often-ignored obvious things into agent memory or project steering markdown or periodic nudges: You must have a hypothesis when you run expensive tests, you must spot check changes first, then start with the most relevant tests only, then move outwards only as necessary to broader labels and only then suites and only then ALL suites in a widening gyre.

But like a falcon ignoring the falconer, the models want to run the everything for anything. So you sigh, you get the model to write a deterministic hook to catch the wrong invocation of the test suite, and you spend weeks refining the rules every time you hit a edge-case, and so it goes. At least you don't have to write the involved regexes by hand, and maybe one day it will be finished..

markkingtoday at 7:48 AM

[flagged]

yuxinkingtoday at 7:43 AM

[flagged]

haukebritoday at 7:27 AM

[flagged]

zhoujinliangtoday at 6:14 AM

[flagged]

runtime_lenstoday at 6:06 AM

[dead]

kestrelquanttoday at 3:36 AM

[flagged]