logoalt Hacker News

lemontheme • today at 6:47 AM • 1 reply • view on HN

I’m still figuring out codemode. I was using it in a prototype, but ended up stripping it out again, after realizing my small local LLM was using more tokens than usual. It was combining tool calls elegantly in code exactly how I hoped it would. The problem was that when any of those embedded tool calls failed (e.g. on parameter validation) the parent code execution tool call also failed. In response, the LLM kept rewriting large parts of the original code block.

Btw, Monty by the pydantic team is a joy to work with if you need a way to securely run unverified code. It’s a simplified Python dialect. You can also use it from JS, iirc.


Replies

hamandcheese • today at 6:55 AM

This somewhat validates a fear I've had (but hadn't tested) of codemode with smaller models. It works fantastically with Opus and Sol but I've always wondered how well it scales down.