logoalt Hacker News

tra3today at 6:00 AM3 repliesview on HN

Love LLMs gonna keep using them. It feels like your suggested approach is expensive, in terms of tokens. I feel (second time I say this) that when I steer the process I get pretty good results vs my coworkers that let the LLMs run away. I do have data on our token usage, not much in terms of quality of the delivery.

I keep thinking about the c compiler implementation that anthropic shared earlier in the year that had all the requirements you mention and arguably wasn’t that great.


Replies

vkazanovtoday at 6:26 AM

Thr thing is that both you and your agent should have a way to verify the solution.

OBVIOUSLY, the compiler experiment was just a cringe pr stunt. But it has a point: everything works better with a good testing loop, and compilers always have one by thr nature of the work they do

cadamsdotcomtoday at 6:32 AM

> expensive, in terms of tokens.

No amount of tokens can come close to my hourly rate.

mrtesthahtoday at 7:17 AM

Do you steer your agents by manually running every single test and linter and reporting the results back to them?