Between this and the Cloudflare post, this is a lot of words and no simple system level picture. Here's what I think is going on:
1. Homoiconicity: Harness mechanics are kludgy and we need proper homoiconicity to uniformly handle code (tool calls) and data ("natural language") as token streams.
2. Actor semantics: for isolation, encapsulation and concurrency.
3. Object capabilities: injected references and no ambient authority. Capabality-based reflection/introspection is a clean way to discover interfaces and affordances.
"Code mode" or whatever is basically rediscovering this by hacking outward from LLM token streams, instead of from system design principles based on decades of computer science. It is the beginning of treating an LLM as a programming-language runtime participant (any takers for eval/apply?) rather than as a text/token generator with the harness as an ad-hoc interpreter.
If existing implementations of code mode don't already support all this, I anticipate they will keep piling on hacks till they get to this point.
----
I think it was Dan Ingalls who said "An operating system is a collection of things that don't fit into a language. There shouldn't be one.". I see the same for harnesses -- they're awkward middle children which fit neither in an LLM nor in the programming environment.
Maybe the answer is to partner LLMs with Common Lisp or Scheme fibers / Spritely Goblins or Erlang BEAM and be done!
It's actually about context management. Calling a MCP tool will load a potentially huge json on your context window, maybe triggering a compaction. Using a subagent for that is less bad, but still expensive (and you wouldn't run every tool call on a subagent)
But if you had a mcp client cli, you can just pipe it to jq or whatever and extract what you need. or chain multiple tool calls into a single one. None of this will make the agent receive the intermediate text passed between those (the agent might want to save the intermediate results into temporary files however)
But shell scripting sucks. Javascript or Python is better suited for handling json and things like that. That's what is being called codemode.
Ok so.. it makes a lot of sense to integrate agents and programming languages. But right now, harnesses are more or less interchangeable and you can hop to another one very quickly. The more coupling between all those moving parts, the hard will lock-in hit. (just picture the mess that is Claude Code having a severe lock in on the ecosystem)
> no simple system level picture
proceed to drop the word homoiconicity... Of course I know what it means without looking up
> Maybe the answer is to partner LLMs with Common Lisp
Autolith (https://autolith.rocks/) and some other CL-based harnesses allow the LLM to modify their own harness within the session.
It's more like this:
Right now when a model wants to call 1 or more tools, there is a fixed json schema to describe the tool calls.
Codemode is like: Why don't we just let the model write a program to call the tools and compose them however it wants? The key is that the programming language exposed to do this (usually javascript) will have APIs available to do some things internal to the harness (like call tools/mcp).
I agree. Harnesses are quickly becoming the next systems programming frontier. The OS analogy is apt and often made--just like browsers have been compared to becoming "platforms" and "OSes". Another analogy that is obvious is threads/processes and agents and sub-agents: all the concurrency, shared resource access, cooperative multitasking, etc problems are emerging again.
It will take some time to shake out, but lots of old ideas are new again. Personally, I'm mining lisp machines, homoiconicity, actor models, and other ideas for my harness designs.