I'm not sure why almost all codemode implementations choose Javascript. I prototyped an agent[1] to use bash as the language for codemode, which in my opinion worked equally well and requires no teaching (there is literally 0 prompt to teach the LLM about codemode. A tool named "bash" is enough to have them know the usage).
> I'm not sure why almost all codemode implementations choose Javascript
Because the models are trained on JavaScript for code mode. You get away with way fewer instructions. They also want to be able to express concurrency and that works very well with the Promise global.
But a big reason is that code mode runs on the harness side so bash is a tricky target in particular.
Almost every model is fully trained on js. That does not need teaching either.
Infact harder to sandbox bash (just-bash or brush based) than it is to js or lua, which has fantastic embedded tooling.
> However bash has one fundamental limitation which is that it can only compose programs that run. And there are some things, which are not programs, but native tools to the LLM and they sort of have to be.
The most obvious example here is `read` or `view_image`. If a multimodal model needs to read an image, it cannot use cat for that because the harness needs to inject the actual image payload into the protocol of the LLM.
Does your prototype overcome the limitations mentioned in the article?