logoalt Hacker News

ed_mercertoday at 12:32 AM2 repliesview on HN

Does this extend to open models like GLM 5.3? This would mean that simply changing the harness to Pi reduces cost in half?


Replies

joshheitzmantoday at 2:10 AM

The provider's middleware also plays a role. I just completed some benchmarks on my bespoke harness and Kilo Code. There's a chart on my LI post here: https://www.linkedin.com/posts/joshheitzman_i-ran-a-fuller-r...

In the case of DeepSeek-V4-Flash-0731 on deepinfra.com there was little difference when both used high reasoning. In the case of that same model on together.ai there was a substantial difference between the two (high reasoning for both again). When using together.ai with Kilo Code the LLM was having a lot of trouble making successful edits. In some cases that meant a lot tries at using the tools and in others it worked around by running scripts. Meanwhile it used the tools from my harness just fine. I've specifically tried to make my tools easy for all of the open weight LLMs to use correctly. That was inspired by getting some errors from Kilo Code at the beginning of the year telling me that the model was having trouble and I should use a smarter model.

roywigginstoday at 2:04 AM

I've found the experience of using Pi with local models feels a lot snappier than both OpenCode or Claude Code.