logoalt Hacker News

gregwebs • today at 2:10 PM • 2 replies • view on HN

I tried the experiment of reviewing vs. not reviewing with frontier models. I consistently found that reviewing by a model with independent context finds important issues when changes are non-trivial- certainly the definition of non-trivial is getting raise as the models get better.

I do have an /implement-simple workflow to skip the planning phase, but even that doesn't skip the review.

Are you doing your own intensive reviews of the model code? Can you share the prompts you are using as I have?

My bar for what models produce without human intervention is much lower defect than what a human would produce. The human interaction is mostly to guide the design and then the review burden is very low. I suspect your bar for what agents produce is lower- you are taking more of the review burden. I also suspect that you are measuring time more than actual cost since your employer is paying and that you are comparing to Sonnet rather than DeepSeek (DeepSeek 4.1 again is 20-40x cheaper than Sonnet). You mention hundreds of millions of tokens (my reviews don't use that much), but even that costs ~$1 on the DeepSeek side.

I think you are taking exactly the right approach at your employer given the cost is free and you only have access to Anthropic models.


Replies

gregwebs • today at 3:14 PM

I saw your response before it was deleted- that you are doing multi agent persona reviews and a very intensive review process. So having fewer review items saves you money.

One thing that I have found is that as the frontier models get better there is less need for agents with specialized personas. I actually don't don't use those anymore- I just use agents that have different models and reasoning levels. I have a generated CODING_STANDARDS.md document and a skill for architecture design and a skill for implementing testing [2] that are referenced by a single reviewer. I do implement a 2-pass review though [3].

I would be interested to know if you have found anything similar as models get better. It seems though that you are sharing a single exploration and then sharing the context across the specialized reviewers to dramatically reduce the cost of your approach. Does this have to be in the harness- that is if you write out the shared context to a file does that increase your costs a lot?

I also wonder how intensively are the models able to test their changes? The number one quality improvement I have found is not review but having the model properly test its code. I have a skill that is helping [4], but I also have to spend time to establish a pattern of testing with tools beyond just unit tests. The testing takes significant effort, and this is again where the cost savings of DeepSeek shine.

  [1] https://github.com/mattpocock/skills/blob/main/skills/engineering/codebase-design/SKILL.md

  [2] https://github.com/gregwebs/skills-sdlc/blob/main/skills/verify/SKILL.md

  [3] https://github.com/mattpocock/skills/blob/main/skills/engineering/code-review/SKILL.md

  [4] https://github.com/gregwebs/skills-sdlc/blob/main/skills/verify/SKILL.md