logoalt Hacker News

gopalrajatoday at 9:45 AM0 repliesview on HN

Yeah — i agree based on my own experience. It was that the loop only works because if you have a rubruc to measure againt. I’ve had similar issues but not in any decompiler mode. Agents can be deceptive in explaining something so well (and I was too lazy to carefully read it verify it myself). Where I got burned was when a beta tester pointed out some citations that didnt exist. So i put a check in place. agent can propose freely, unconfirmed stuff fails closed, my rubric checks get the veto. Same shape as what you’re describing. The agent can refine; the rubric metrics is what keeps it from just becoming fluent drift. Without that human in the loop, autonomous improvement is mostly relies on luck. On the other hand, I am seeing more and more the newer model's reasoning and having another llm eval (provided that llm has rubrics to measure against) does lessen the falseness.