The kernel you share with a stranger
A container is a process in a costume; it shares the host kernel with everything else on the box, so its isolation is the kernel's bug count. Author and lifetime pick the boundary.
A container is a process wearing a costume. It has its own view of the filesystem, its own process table, its own network namespace, and it shares the kernel with every other process on the machine, including the ones that belong to someone else. The isolation a container gives you is therefore exactly the isolation the kernel gives between processes, which is to say it is as strong as the kernel's bug count is low, and the kernel is a large program that gains and loses bugs every week. Whether that matters depends on two questions, and once they are answered the boundary picks itself: who wrote the code, and how long will it run?
The stranger's code test
The first question is about trust, and it has four honest answers. You wrote the code, and you can read it. A dependency wrote it, someone else whose work you vetted when you added it and whose updates you review, more or less. A customer wrote it, a person you have a contract with and no visibility into. Or a language model wrote it, which is the newest answer and the one this post exists for, because model-written code is code whose author cannot be held to a contract, cannot be asked what they meant, and will produce something different next time.
The second question is about exposure. Code that runs for one request and exits has a small window to find a kernel bug and use it. Code that runs for a session, minutes to hours, has a larger one. Code that runs indefinitely, a long-lived worker or an agent that keeps a loop going, has all the time there is, and an attacker who has all the time there is will eventually find the bug that the last kernel patch did not cover.
Read the grid from the top left. Your own code, however long it runs, is a process; the kernel's process isolation is what it was built for and you are the only author. A vetted dependency gets a container once it runs for any length of time, because a container adds the filesystem and network fences that limit what a bug in the dependency can reach, and those fences cost nothing. A customer's code, for a single request, is also a container; for a session it wants a userspace kernel such as gVisor, which intercepts the system calls and answers most of them without touching the host kernel, so the attack surface shrinks to the calls the userspace kernel passes through; and for anything indefinite it wants a microVM, a separate guest kernel behind a hardware virtualisation boundary.
A language model's code starts one column further right than a customer's, because there is no contract and no author to call. Even for one request it gets a userspace kernel. For a session or longer, it gets its own kernel. The bottom-right cell, model-written code running indefinitely, is the corner where the shared kernel is the whole risk and nothing short of not sharing it is acceptable.
The excuse Firecracker removed
The usual objection to microVMs is cost: a virtual machine takes seconds to boot and hundreds of megabytes to exist, so you cannot give one to every request. That objection describes a general-purpose hypervisor, and Firecracker was built to remove it. Its specification commits to at most 125 milliseconds from the API call that starts a microVM to the guest's init process, and at most 5 mebibytes of memory overhead for the virtual machine manager's threads, measured on bare-metal hosts; the paper that introduced it reports starting 150 microVMs a second on a single host. A boundary that costs an eighth of a second and five mebibytes is a boundary you can afford per session and, for many workloads, per request.
The reason those numbers are possible is the same reason the boundary is strong: Firecracker implements a deliberately small device model, a few virtual devices and nothing else, so there is little for a guest to attack and little to initialise. A general-purpose hypervisor emulates a whole machine. A microVM emulates just enough of one to run a kernel, and the missing parts are missing attack surface.
What the engine class buys and what it does not
A 2026 comparative study of sandboxes for model-written code, Andronchik and Lokhmakov, looked at five products across the three engine classes, microVMs, userspace kernels and containers, and reached three conclusions that fit the grid. The engine classes separate cleanly on every architectural axis, which is the justification for choosing by class first. Products within a class do not separate, so the choice between two microVM products is about operations rather than architecture. And the property that dominated real outcomes was patch latency: engine-side patches for coordinated disclosures landed in roughly zero days, while downstream lag, the time before a product actually shipped the patched engine, ranged from zero days to more than 471, to opaque, to never.
That third finding is the one to carry into procurement. The grid tells you which class of boundary a cell needs. It does not tell you that a product in that class is patched, and a microVM running a guest kernel that is 471 days behind is a separate kernel with 471 days of known holes. The two questions that pick the boundary get a third, operational one: how quickly does the product you run ship the engine's fixes, and can you see the answer?
What I would require before letting a model's code run
None of the products I have shipped executes model-written code, so what follows is a requirement rather than a report. Before a product of mine ran code that a model produced, I would want three things written down. The cell in the grid the workload occupies, which for an agent that keeps state between steps is the bottom row and rarely the left column. The boundary that cell demands, which for the bottom row is a userspace kernel at minimum and a microVM for anything that lasts, with the microVM's cost now small enough that the argument for the weaker option is thin. And the patch policy of whatever runs the boundary, with a number of days attached, because the study's finding is that the number varies by three orders of magnitude between products that look the same on a feature list.
There is a fourth thing worth writing down, which is what the boundary is for. A sandbox is not a substitute for limiting what the code can reach: a microVM with the production database's credentials mounted inside it is a separate kernel with a straight road to the data. The boundary protects the host from the code; the credentials, network policy and filesystem mounts protect everything else from it, and the grid says nothing about those because they are the same in every cell.
The kernel is the thing you share with a stranger when you run their code in a container, and the stranger, increasingly, is a model. The costume is fine for your own code. For theirs, give them their own kernel. It costs an eighth of a second.
Get new posts by email
Occasional essays on engineering, AI, and building for the people technology leaves behind.
Subscribe with RSS