> However bash has one fundamental limitation which is that it can only compose programs that run. And there are some things, which are not programs, but native tools to the LLM and they sort of have to be.
The most obvious example here is `read` or `view_image`. If a multimodal model needs to read an image, it cannot use cat for that because the harness needs to inject the actual image payload into the protocol of the LLM.
Does your prototype overcome the limitations mentioned in the article?
Surely LLMs understand how to decide base64, but if that’s not possible you could convert the image to ascii and then send that in
There's no limitation, just have bash commands `read` or `view_image` which communicate back to agent.