Do any of these sandboxing solutions have a dynamic component to them that lets you grant permissions, starting with a minimal sandbox and asynchronously adding permissions as they become necessary? Harnesses try to do this when accessing non-project folders, but it's not always strictly enforced and generally not revocable. Harnesses also block agent execution until a decision is made, which requires constant monitoring to ensure progress can happen when the agent could easily proceed with an alternative method right away.
I like the idea of a minimal sandbox that protects against accidental `rm -rf` and against personal data leakage, but such a setup then often gets in the way of the specific task to be done. Ideally the sandbox would be able to aggregate blocked accesses and then expose them in an external TUI dashboard, where I can then enable access (without blocking any running agent on this, since that's prone to "press okay" fatigue).
Does anything close to this exist yet?
Seems aws also announced a sandbox solution
I used agent over 1 year and basically always give codex full permission on each thread, do not get issue so far
Why we need this layer of complexity? Or its mainly for big company that need control ?
350,000 of mostly Rust SLOC [1] ... And the upstream sandboxes aren't even vendored!
I'd be way more confident building upon something I can grasp and understand. [2]
[1]: https://ghloc.dev/microsoft/mxc [2]: https://github.com/sandbox-utils/sandbox-run
Microsoft stole my idea :) (joking obviously, everyone and their mom is making sandboxes) https://github.com/pprotas/slopbox
Why is everyone making their own code execution agent runtime engines I have an entire project built on top of openshell already, why not first come up with a sandboxing policy design, like unix did, and then build on top of that.
Currently all project do tend to agree on what and how they work but certain things being different makes porting tedius, if all of them have a bare minimum subset common amongst them it would be much easier to switch, and validate security surface area.
I feel like there are more vulnerabilities in this vibe coded slop sandboxes, and it's more likely everyone one of us trusting them to build projects around them will shoot our foot off once a cve is hit in one that's common in all of them but since they are all slop copies someone will have to figure out how they apply to all others and then manually fix it properly, and if one of them makes a CVE public it will leave dozens of these runtimes open to exploits.
I wish the best to my future self with regards to security I feel like we are completely screwed. Since we can no longer depend on upstream for security.
Would this allow a sandboxed container on windows to still run commands in wsl?
They have a sandbox escape in there. Likewise capability ordering is wrong. Exactly what you should expect from Microsoft.
Since this apparently wraps bubblewrap (another incompetent action on behalf of Microslop), did a quick sweep of that codebase too. Setuid is wrong, capability dropping is wrong, bubblewrap does _not_ protect against compromised/vulnerable kernels (and it should fyi), wrote up a full bubblewrap/mxc sandbox escape too.
[flagged]
Asked Grok the difference between mxc and flatpak:
"So MXC is a cross-platform “what may this workload touch?” layer aimed at agents. Flatpak is a Linux app format whose sandbox happens to share a backend with MXC on Linux."
Saw the M stands for Microsoft and immediately closed the tab. 1 it’s unnecessary - communicates nothing but look-at-me branding. 2 toxic company.
This looks pretty decent actually. Sure, you could consider it a frontend/SDK for bubblewrap/seatbelt/processcontainer; but setting em up consistently is far from trivial; and hand rolling is a really bad idea (speaking from experience).
I like the ‘learning’ mode for figuring out what perms/config a runtime needs, the MIT license, the clear optional telemetry disclosures, and somewhat light and still readable documentation.
Regardless of your views on Microsoft, this looks quite useful; serves a clear purpose, and from a quick glance, looks like a high quality project even if it’s just the first version.