I'm a bit bumped when I first saw pi is moving from bash to codemode and also adding MCP, since I thought that loses the purity and simplicity of the "use bash for everything" philosophy. However, after reading more about it, I realized codemode is just a slightly enhanced version of bash: more complex, sure, but likely more robust, secure, and efficient. For those who, like me, don't get the point of this change, here is how I understand it.
The first tool execution runtime in harnesses are direct tool calls with JSON or XML, such as the Read and Edit tools. As an escape hatch, we have Bash tool that allows arbitrary code execution on the host running the agent. The downsides of using bash (on the host) as the main tool execution runtime are:
- Syntax and obvious errors only surface at runtime
- Unergonomic orchestration of parallel and background tasks
- Verbose command output cluttering context
- Dependent on the host environment, packages versions, etc.
- No security measures by default.
To me the last point is the biggest inherent weakness, usually mitigated by creating a dedicated unprivileged user or running bash in a sandbox.
Note that direct tool calling is kind of the polar opposite on these points: syntax errors are caught early, orchestration can be done with some wrapping tools, command output is controlled, and most importantly they are more sandboxed. On the flip side, they obviously have way less power, necessitating Bash tool in the first place.
Codemode is the middle ground between these two extremes. It actually can be derived simply by one idea: what if we replace Bash by another language that can be checked for obvious errors, i.e. type checked?
Everything else falls out from there:
- Any language would do, but I think TypeScript fits the balance between safety, speed, conciseness, and popularity in training data.
- If we use TypeScript, might as well run it in a sandbox as JS runtimes have been designed with this in mind for 20 years
- Orchestration comes for free from the JS runtime. It's not more powerful, just more ergonomic.
- Since the tools are controlled by the harness and not dependent on the host, cloud agent becomes easier.
- With this in place, MCP are not very different from a tool provided to this sandboxed runtime.
Overall I find the benefits compelling enough, but we'll see if the heavily-RLed models these days will use it effectively.