MXC - a sandboxed code execution system

(github.com)

98 points | by nreece 9 hours ago

23 comments

  • dannyw 6 hours ago
    This looks pretty decent actually. Sure, you could consider it a frontend/SDK for bubblewrap/seatbelt/processcontainer; but setting em up consistently is far from trivial; and hand rolling is a really bad idea (speaking from experience).

    I like the ‘learning’ mode for figuring out what perms/config a runtime needs, the MIT license, the clear optional telemetry disclosures, and somewhat light and still readable documentation.

    Regardless of your views on Microsoft, this looks quite useful; serves a clear purpose, and from a quick glance, looks like a high quality project even if it’s just the first version.

    • gandreani 2 hours ago
      I agree! I can't comment on the design of the SDKs but from the looks of it this is a great option to integrate into agent harnesses/pipelines.

      The "audit" and "debug" mode are specially useful. I use bubblewrap and using a new harness with it is usually a couple of rounds of wack-a-mole with strace to figure out all the harnesses dependencies.

    • LeBit 4 hours ago
      It does look nice and easy.

      Not sure I will give up on smolvm though.

      I have scripts to launch an instance per task. Nono wraps my coding agent.

      I provide the git clone for the specific task.

      It works really well.

  • neobrain 3 hours ago
    Do any of these sandboxing solutions have a dynamic component to them that lets you grant permissions, starting with a minimal sandbox and asynchronously adding permissions as they become necessary? Harnesses try to do this when accessing non-project folders, but it's not always strictly enforced and generally not revocable. Harnesses also block agent execution until a decision is made, which requires constant monitoring to ensure progress can happen when the agent could easily proceed with an alternative method right away.

    I like the idea of a minimal sandbox that protects against accidental `rm -rf` and against personal data leakage, but such a setup then often gets in the way of the specific task to be done. Ideally the sandbox would be able to aggregate blocked accesses and then expose them in an external TUI dashboard, where I can then enable access (without blocking any running agent on this, since that's prone to "press okay" fatigue).

    Does anything close to this exist yet?

    • __MatrixMan__ 1 minute ago
      If only we had a decent capability based OS. Then all processes would have the property you're after. If you want it to be able to write to a disk, dependency inject a disk writing capability. Need to talk to a remote host, provide a handle for just that host.

      Don't want these capabilities? Do nothing, that's the default state.

    • 0kk33 3 hours ago
      If I understand you correctly https://nono.sh/ might go into that direction. It can add permission after the sandboxed command is terminated based on which blocks occurred. Its not life though as you seem to describe
      • Neywiny 13 minutes ago
        I couldn't get cline TUI to run in it though. Apparently node does fancy temp folder things and it just wouldn't go
      • neobrain 1 hour ago
        Very interesting concept to make permissions specific to individual shell commands though. Certainly good that people are experimenting with these approaches, hopefully ideas will eventually converge so we don't need to know like a 100 different sandbox projects :)
    • ghm2180 2 hours ago
      I think this may be a false choice because of the work pattern you might be used to. Assuming here, If you work in interactive sessions where it's open ended there is a boundary where you have done enough research/prototyping and you need to move to implementation and the permission scope has to change now. I think the realization that you might have is that if you're doing this then it's probably best to separate the automated AFK part from your initial research part.
      • neobrain 2 hours ago
        Even for pure research/prototyping, you quickly run into the problem that your sandbox is either prohibitively minimal or overly permissive. Depending on the exact task, you may need GPU access, Docker/Nix socket access, ability to ptrace processes (gdb), run webfetches, etc. If I define a "research" sandbox profile to allow all of these, I might as well not have any sandbox at all.
    • ghm2180 3 hours ago
      I have faced this exact dilemma as well. It's not always clear what permissions are needed in advance for my pi sessions and it's child sub sessions. A simple example is when child sessions do a task they locally want to fire random docker commands to learn the state of my local docker devstack.
    • agentdev001 2 hours ago
      Take a look at Nvidia OpenShell, and their dev blogs on it. That seems like what you're looking for.
      • neobrain 1 hour ago
        Sounds OpenShell locks filesystem access on sandbox creation - arguably the most important isolation feature, at least for my use cases. Architecturally it looks right though!
  • simonw 1 hour ago
    This is promising, but there's one feature that's missing that I really care about: fine-grained networking.

    They have this for Windows and Linux, but it's sadly missing for macOS - see the support table here: https://github.com/microsoft/mxc/blob/main/docs/backends/sea...

    Things macOS is missing include "Allow/deny by hostname" and "Allow/deny by IP, CIDR, port, or protocol".

    The rest all looks great, and if you are on Linux or Windows those restrictions don't apply.

    I guess this is the universal challenge of building an abstraction layer over multiple different technologies.

    • gregwebs 1 hour ago
      Fine grained network policies is supported by microsandbox- a project that has already been working hard at building an abstraction layer over multiple different technologies. Microsandbox (on unix) builds on top of libkrun (a VM abstraction layer for unix). I am building a convenient runner on top of microsandbox: https://github.com/runcontain/runcontain (undergoing a rename right now). The best thing Microsoft could contribute right now would be great technology for light-weight containment on Windows.
    • dannyw 1 hour ago
      They could bundle in a HTTP proxy (enforcing similar rules) perhaps. It takes a bit of reading to dig-through the Claude speak, but "Egress confinement is enforced; using the proxy is cooperative" simply means that there's no network egress, except through the proxy.

      Of course, that only limits HTTP; and not other forms of network requests.

    • simonw 1 hour ago
      ... interestingly, Anthropic's SRT is built on the same macOS primitives and DOES support the network configuration I'm looking for:

      https://github.com/anthropics/sandbox-runtime/tree/main#as-a...

        const config: SandboxRuntimeConfig = {
          network: {
            allowedDomains: ['example.com', 'api.github.com'],
            deniedDomains: [],
          },
          filesystem: {
            denyRead: ['~/.ssh'],
            allowWrite: ['.', '/tmp'],
            denyWrite: ['.env'],
          },
        }
  • epage 2 hours ago
    Been looking at sandboxing, both low level and higher level like this.

    The API for their Rust mxc-sdk looks nice but

    - their "sdk" has binaries and the build script has logic for them

    - their build scripts do windows-exclusive work on all platforms

    - not putting some of the backends behind features causes more build script work (and that work will break on future Cargo versions)

    - at least some of the remaíning build script work doesn't need to be a build script

    - it seems pretty dependency heavy

  • kernc 5 hours ago
    350,000 of mostly Rust SLOC [1] ... And the upstream sandboxes aren't even vendored!

    I'd be way more confident building upon something I can grasp and understand. [2]

    [1]: https://ghloc.dev/microsoft/mxc [2]: https://github.com/sandbox-utils/sandbox-run

    • dannyw 5 hours ago
      If you look around the files, I think at least half is comments or unit tests, e.g.

      https://ghloc.dev/microsoft/mxc?branch=main&locsPath=%5B%22s...

      That site thinks this file has 2.9k sloc and doesn't seem to parse rust comments. In reality, there's only 1,465 sloc; and 635 loc of tests.

      Definitely nowhere near 350k sloc.

      --

      As for your sandbox run: it's a single-contributor project, seems to have only have basic smoke tests, and has a few major/critical security issues:

      * _generate_seccomp_filter compares newline-deliminated syscalls, against a multi-line blocklist, meaning the entire function doesn't block anything and is essentially a no-op.

      * Main script invokes working directory's .env as shellcode, before switching into restricted filesystems and dropping capabilities. Attacker-controlled .env can run shellcode with full privileges.

      * Lots of race conditions which I haven't verified, but doesn't really matter.

      I'd make PRs, but I don't think it's a good idea to try and DIY a sandboxing system in bash with minimal SLOC as the target in the first place. I'm also slightly concerned that most of your comments on HN seem to be promoting this repo?

      • kernc 2 hours ago
        > at least half is comments or unit tests

        Thanks, I see there's a slight (~50%?) overestimation there, but then again, even unit tests and comments in a target programming language count as syntactically correct code that needs to be evaluated and reasoned upon. I'm not that familiar with Rust's runtime introspection features, but in languages like Python, even the comments can directly affect code (e.g. `Foo.__doc__ = Bar.__doc__ + SOME_ANNEX`).

        > _generate_seccomp_filter ... the entire function doesn't block anything

        Many thanks! I've applied a fix—it's a single line added. The missing test is pending a runnable that invokes one of the forbidden syscalls. As I have no qualms about force-pushing around a repo that nobody forks, happy to credit you(r LLM) proper!

        > Main script invokes working directory's .env as shellcode

        The sandboxed process can't overwrite existing .env files [1], but it could create a new $PWD/.env file, hoping to "escape" at next sandbox execution. That's a valid concern I'll have to think about some more.

        [1]: https://github.com/sandbox-utils/sandbox-run/blob/c97d065184...

        > Lots of race conditions

        I sometimes experience "Slirp not ready in time" [2], but it's due to a so far unexplained upstream issue [3]. I you have time/tokens to spare, I'd appreciate those PRs and further similar feedback!

        [2]: https://github.com/sandbox-utils/sandbox-run/blob/c97d065184... [3]: https://github.com/rootless-containers/slirp4netns/issues/35...

        Don't know whether it's a good idea. It sure has got its issues. But even as the SLOC count and the number of bugs metrics are proved correlated in literature [4], min SLOC is not the primary target—a reasonably graspable and stable composition of few dependencies is. Whereas overreliance on third parties nowadays often ends with a rug pull one way or another. We simply can't count on this "MXC" (...) to be maintainable/non-archived even a year from now, just when I'd get it all properly integrated and set up.

        [4]: https://softwareengineering.stackexchange.com/questions/1856...

        > slightly concerned

        Oh, I certainly wouldn't like to limit myself to promoting just this repo! ^D^ HN is a good venue, lots of smart people around! I see everyone shilling their own sh** all the time. Often in green usernames. :shrug:

      • IshKebab 4 hours ago
        And it's 600 lines of dense Bash. I trust 350k lines of Rust way more than that!
        • kernc 2 hours ago
          600 lines of dense POSIX Shell—in some respects that's even worse!

          You would not be aware of the amount of trust you are putting into that.

    • its-summertime 3 hours ago
      I feel a better metric is `lines changed / time` as that affects what will be audited as time goes on

      That being said, a month of MXC has more line changes than 2-3 years of runc

  • eminence32 2 hours ago
    A little off topic, maybe, but I've been having great luck with wasmtime and wasm32-wasip3 for writing sandboxed plugins. The tooling is pretty nice when you write plugins in rust, but I don't know what it looks like for other languages right now.

    wasip3 is not stable yet, but it has a lot of nice changes (compared to wasip2) for integrating with async code

  • pprotas 4 hours ago
    Microsoft stole my idea :) (joking obviously, everyone and their mom is making sandboxes) https://github.com/pprotas/slopbox
  • wild_pointer 2 hours ago
    What's also interesting is that they added the Experimental_CreateProcessInSandbox API to Windows, like, last month.

    https://learn.microsoft.com/en-us/windows/win32/secauthz/cre...

    Funny, now that it's documented, the API name will be stuck with this name forever.

    • mrpippy 8 minutes ago
      The docs are pretty clear that it’s experimental, and it’s not even exported from a DLL. It’ll be hilarious if some application depends on it and they have to keep it though
  • mintflow 3 hours ago
    Seems aws also announced a sandbox solution

    I used agent over 1 year and basically always give codex full permission on each thread, do not get issue so far

    Why we need this layer of complexity? Or its mainly for big company that need control ?

    • pprotas 2 hours ago
      The main usecase for an average developer is preventing confused agents making mistakes like removing sensitive folders, resetting git branches or using API tokens they shouldn't be using
    • dannyw 1 hour ago
      A ~month ago, auto-review (rightfully) blocked a rm that would've nuked my home directory, due to shell mangling (amongst other issues).
    • joshuanapoli 3 hours ago
      If you have a custom agent in a product, then it needs isolation to be sure to protect the customer data.
      • stingraycharles 2 hours ago
        People are running custom agents that are able to run custom code in production just like that ?

        I’d personally opt for SELinux in such cases

  • minraws 6 hours ago
    Why is everyone making their own code execution agent runtime engines I have an entire project built on top of openshell already, why not first come up with a sandboxing policy design, like unix did, and then build on top of that.

    Currently all project do tend to agree on what and how they work but certain things being different makes porting tedius, if all of them have a bare minimum subset common amongst them it would be much easier to switch, and validate security surface area.

    I feel like there are more vulnerabilities in this vibe coded slop sandboxes, and it's more likely everyone one of us trusting them to build projects around them will shoot our foot off once a cve is hit in one that's common in all of them but since they are all slop copies someone will have to figure out how they apply to all others and then manually fix it properly, and if one of them makes a CVE public it will leave dozens of these runtimes open to exploits.

    I wish the best to my future self with regards to security I feel like we are completely screwed. Since we can no longer depend on upstream for security.

    • lifeisloving 4 hours ago
      I saw a tweet that said:

      "Im really fkin worried we're all building the same thing"

      Everyone has been building a harness/sandbox the last 6 months. Ive seen dozens and dozens shared in discords.

      Even companies are totally stuck focused on the same paradigms.

      The previous iteration of this was RAG/Chat interfaces. See PewDiePie's project. Last month it was briefly everyone building the same classifier.

      Peter Thiel, gave a lecture about this same phenomenon 15 years ago likely because he observed the same things going on during other hype cycles. Everyone building the same things. Its called something like "Dont build the obvious thing"

      This is why Im moving towards hardware for personal projects, it forces me to be much more creative and think outside the "How can I make something AI adjacent/powered" trap thats so easy to fall into in pure software right now.

    • torginus 5 hours ago
      The problem with OS level sandboxes, and the reason why WebAssembly's being explored in this space (and Electron is so popular), is that relying on OS/hardware features means your TAM shrinks to a fraction of total, and it's historically well known you set yourself up to lose.

      History is littered with tons of super cool OS features that didn't manage to gather enough market share and ended up as cool futures, and fodder for 'we invented the future 20 years ago' style articles.

    • hobofan 6 hours ago
      Different use-cases have different requirements.

      e.g. this one puts multi-platform support as a high requirement, a requirement that OpenShell doesn't fulfil (and likely won't given it's architecture/goals).

    • fg137 3 hours ago
      My guess is that Microsoft thinks this can be deployed with standard, company-wide policy across platforms (mostly) with their IT management tools which poses a unique advantage.

      In reality, however, knowing how much difference there is between OSes, how tricky it is to configure these things to make them actually useful, and how bad Microsoft products are, I'm not enthusiastic about this project -- there are so many others on the market already, and I'll wait to see if this gains traction.

      (Notice that on MacOS it only supports seatbelt? That's not nearly the same as microvm.)

    • rock_artist 5 hours ago
      That’s exactly it. There should be some permission logic for delegating.

      But as always, there are rivals trying to set their tone on what’s the standard. We all wish there was one unified agreed concept that will work but I guess the most common one will eventually survive.

      Just as Microsoft in a sense embraces Linux with WSL and also Apple has their virtualization framework.

      I hope we’ll eventually get unified model management system to include also permissions designed properly

    • booster-rooster 5 hours ago
      [flagged]
  • chneu 7 hours ago
    Right you are, Ken!
  • superxpro12 24 minutes ago
    Right you are, ken!

    But seriously, we need docker for models like years ago. I dont want these things running with the ability to run rm -rf /

    its should be treated no different than wget | sh

  • fassssst 1 hour ago
    Codex uses this on Windows now.
  • arj 5 hours ago
    Would this allow a sandboxed container on windows to still run commands in wsl?
  • ranger_danger 1 hour ago
    How does this compare to Sandboxie?

    https://github.com/sandboxie-plus/Sandboxie

  • rfgplk 4 hours ago
    They have a sandbox escape in there. Likewise capability ordering is wrong. Exactly what you should expect from Microsoft.

    Since this apparently wraps bubblewrap (another incompetent action on behalf of Microslop), did a quick sweep of that codebase too. Setuid is wrong, capability dropping is wrong, bubblewrap does _not_ protect against compromised/vulnerable kernels (and it should fyi), wrote up a full bubblewrap/mxc sandbox escape too.

    • dannyw 1 hour ago
    • zbentley 2 hours ago
      Could you link to issue reports to back that up (or maybe file them if these are novel findings)?

      If that’s too much of an ask, at least reference the code you found for things like “mxc sandbox escape” or “bubblewrap setuid is wrong”. Those claims require evidence.

  • mouldloft 21 minutes ago
    [flagged]
  • arbor-group 1 hour ago
    [flagged]
  • singularityisne 3 hours ago
    [flagged]
  • dcmatt 1 hour ago
    [flagged]
  • t_privos 3 hours ago
    [flagged]
  • smitty1e 5 hours ago
    Asked Grok the difference between mxc and flatpak:

    "So MXC is a cross-platform “what may this workload touch?” layer aimed at agents. Flatpak is a Linux app format whose sandbox happens to share a backend with MXC on Linux."

  • zenapollo 3 hours ago
    Saw the M stands for Microsoft and immediately closed the tab. 1 it’s unnecessary - communicates nothing but look-at-me branding. 2 toxic company.