Infrastructure Cost Structures That Make Dollar-per-Agent Pricing Possible
Firecracker micro-VMs and snapshot-restore scheduling enable cheap per-agent pricing.

Dollar-per-agent pricing works because of a specific stack of infrastructure decisions, not because of a billing trick. The thesis of this piece is simple: Firecracker micro-VMs on bare metal, combined with snapshot-restore scheduling, let a provider charge for storage during idle time and for compute only when an agent is actually working. Every layer of that stack exists to solve one problem, which is that agent workloads spend most of their life waiting, not computing.
Why agent workloads are structurally different from the compute they run on
An AI agent spends most of its session doing nothing at all, in the computational sense. It is waiting: for a language model to return a response, for a tool call to finish, for a database query to come back with a result. During all of that waiting, the agent is blocked on I/O, sitting idle while some other system does the work it asked for.
Traditional compute billing was not built for this pattern. A developer renting a server pays for vCPUs for the full duration of a session. That means paying for CPU-hours regardless of how much of that time the CPU actually spends computing. For a workload where real computation might fill a small fraction of total wall-clock time, that billing model charges for hours of silence. The gap between what gets billed and what gets used is not a rounding error. If every idle minute costs the same as every active minute, there is no way to offer a cheap, dedicated agent per user and still make the economics work; rethinking the compute layer underneath it is what makes per-agent pricing at low cost possible.
Why containers cannot solve agent isolation
Agent workloads add a second constraint on top of the idle-time problem that rules out the cheapest fix: agents routinely run code they did not write themselves, outputs from tool calls, scripts supplied by users, routines generated on the fly by a language model. Running that code safely next to other tenants' workloads requires isolation at the hardware level, and that requirement does not bend for the sake of convenience.
Containers, the default tool for cheap multi-tenant compute, separate processes using Linux namespaces and cgroups, but every container on a host still shares that host's kernel. A vulnerability in the kernel can be reached from inside any one container and used to touch the host itself, along with every other tenant running on it. For workloads where the code running inside the sandbox is fixed and trusted, that risk is often judged acceptable. For an agent that writes its own Python scripts, installs packages on demand, and manipulates file descriptors, it is not. The shared kernel surface becomes a liability the moment the code running against it is untrusted, and the blast radius of one bad script expands from a single disposable container to the entire host.
This is why the major cloud providers have moved away from plain containers for their own multi-tenant runtimes. AWS Lambda and Fargate run on hardware-virtualized micro-VMs. Google Cloud Run and GKE use gVisor to sandbox system calls between the container and the kernel. Both are responses to the same fact: container boundaries are not strong enough for code a provider cannot vouch for.
Full virtual machines solve the isolation gap but introduce a cost of their own. A full VM can take tens of seconds to boot and carries memory overhead in the hundreds of megabytes. Allocating a dedicated VM per session, per user, at any real scale, is not a security problem so much as a cost problem: the overhead per instance makes the whole idea financially unworkable long before isolation becomes the limiting factor.
Firecracker and Bare Metal as Substrate
Firecracker was built at AWS to close exactly this gap. Lambda needed to run millions of functions for strangers on the internet, which meant two constraints at once: isolation strong enough for code from parties AWS had no reason to trust, and overhead low enough to make running at that volume affordable. Traditional VMs met the first constraint and failed the second. Containers met the second and failed the first. Firecracker is the design that satisfies both: a micro-VM monitor built specifically for secure, multi-tenant, minimal-overhead execution of container and function workloads.
A Firecracker micro-VM boots a minimal Linux guest in about 125 milliseconds, with memory overhead of around 5 MB per instance. It gets there by using a deliberately small device model: virtio net, virtio block, virtio vsock, a serial console, a minimal keyboard controller, and a metadata service, nothing more. A smaller device model means a smaller attack surface and a smaller, simpler virtual machine monitor to secure. Each micro-VM runs as its own Firecracker process on the host, with its own vCPU threads and its own memory, backed by four layered security boundaries: KVM hardware virtualization, seccomp filters applied per thread, cgroups and namespaces for resource isolation, and a jailer process that drops privileges before Firecracker itself ever executes. A single host running Firecracker can produce up to 150 new micro-VMs per second, and that density is what makes it viable to allocate one micro-VM per user, not just per large customer.
None of that performance is guaranteed by Firecracker alone. It depends on what Firecracker is running on. On bare metal, Firecracker talks directly to the host's KVM. Run it inside a VM instead: every I/O request has to pass through the guest's virtio driver and then through the outer hypervisor before it ever reaches physical hardware. That gap matters most for exactly the workloads this piece has been describing: agents that spend most of their time blocked on I/O will feel every extra hop in the I/O path, because I/O is nearly all they do. Bare metal gives maximum micro-VM density per host, the lowest tail latency, and direct access to devices like local NVMe storage and SR-IOV network interfaces. Nested virtualization adds a "nesting tax" on every I/O operation, and that tax compounds precisely where agent workloads live. KVM itself is non-negotiable. Firecracker runs on x86_64 and aarch64 Linux hosts with KVM, with no emulation fallback, so the hosting substrate underneath Firecracker decides whether its performance numbers are real or theoretical.
How snapshot-restore scheduling converts idle time into a non-event
Firecracker's speed and density set up snapshot-restore scheduling, which is what actually produces dollar-per-agent pricing. The moment an agent blocks on I/O, waiting on a model response or a tool call, its micro-VM can be frozen to disk in full. The CPU core it was using is freed immediately for another workload. When the I/O completes and the agent needs to resume, it wakes from that snapshot in under a second, with its memory, disk, and kernel state exactly as they were. There is no re-initialization, no re-authentication, no lost context. The agent simply continues.
While an agent sits in that frozen state, the only resource it consumes is storage, on its own persistent disk sitting on cheap media. That is the entire mechanism behind the pricing model: if an agent is only active for a small share of any given hour, a snapshot-restore scheduler charges for that small share of compute time plus continuous, cheap storage, instead of a full hour of reserved vCPU regardless of use.
The economics scale further because of how memory gets allocated across a fleet of these micro-VMs. Picture a hotel that sells more rooms than it has, on the reasonable bet that not every booked guest shows up on the same night. A host running Firecracker can configure far more micro-VMs than its physical RAM would support if all of them were active simultaneously, because Firecracker only allocates memory pages that are actually in use, and because most micro-VMs on that host are snapshotted, consuming no RAM at all, at any given moment. That oversubscription is what multiplies the whole fleet's economics, turning a fixed amount of physical memory into capacity for a much larger number of mostly-idle agents.
Persistent state as an architectural requirement under snapshot-restore scheduling
If an agent's state cannot survive being paused, none of the scheduling model above works. Durable storage for its files, durable credentials, and a durable record of what it was doing are not conveniences layered on top of agent infrastructure. They are what makes the pricing model function. If an agent loses its installed packages, its browser session, or its in-progress task state every time it goes idle, it cannot be rescheduled cheaply. It would have to rebuild its environment from nothing on every wake, and that reconstruction cost erases the entire savings that snapshot-restore was supposed to deliver. Idle time would go back to being expensive, just in a different ledger.
The failure mode this guards against is concrete. An agent can hit a rate limit mid-task, time out waiting on an external system, or pause to wait for a human to approve its next step. In any of those cases, it needs to resume exactly where it left off, without redoing work that already finished. The state of that task has to live somewhere durable, not only in the memory of a process that might not exist anymore by the time the agent wakes. Each agent needs its own persistent disk for this to hold, not a volume shared across agents and not a network filesystem mounted into a shared container. The structure that falls out of this requirement, one agent, one micro-VM, one disk, happens to be the same structure that correctly isolates one user's data from another's, a security property that cost pressure produced on its own.
How the per-user agent model becomes economically viable at fleet scale
Put the three pieces together, Firecracker's density, snapshot-restore scheduling, and per-agent persistent disks, and dedicating one micro-VM to every single user becomes something a provider can actually afford to offer. The cost of an idle agent collapses down to the cost of its storage. The cost of an active agent scales with how much it actually runs, not with how many agents exist in the fleet.
A fleet built this way gives every user a genuinely isolated micro-VM, with its own disk, its own credentials, and its own session history, rather than a stateless function that forgets everything between calls. At any given moment, the overwhelming majority of agents across that fleet are idle: between tasks, waiting on their user, asleep overnight. The fleet's active compute bill at that moment is set by the small fraction of agents currently doing work, not by the total number deployed. That is what lets a provider quote a flat monthly rate per agent instead of an hourly compute rate that would be too expensive for most users to justify.
Multi-tenant isolation at this scale tends to follow one of two patterns, and both depend on the same underlying guarantee. In a silo model, each tenant gets a dedicated runtime with its own execution role and its own isolated storage. In a pool model, agent logic is shared across tenants, but state is still partitioned by tenant context. Either pattern only holds if the boundary between tenants is a real hardware boundary enforced by the micro-VM, not a software namespace that a determined or buggy process might cross. Cost control belongs at this same layer: caps on compute per agent and per tenant need to live in the infrastructure itself. A runaway loop or an unexpected spike in one agent's workload gets caught there before it shows up on anyone's bill.
What this architecture demands from the infrastructure layer
Dollar-per-agent pricing comes out of a specific stack; it is not a billing policy a provider can bolt onto any compute backend. If you are evaluating infrastructure for agent deployment, you can check for that stack directly. The requirements are architectural: direct access to KVM, whether that comes from bare metal or from nested virtualization with genuinely verified device access; Firecracker or an equivalent micro-VM monitor; a snapshot-restore scheduler with wake latency under a second; and a persistent disk attached to each agent that survives repeated snapshot cycles intact.
Nested virtualization is not a substitute for bare metal here. Because agent workloads spend so much of their time in I/O, the nesting tax from running a hypervisor inside a hypervisor compounds on every tool call and every file operation, and a platform that runs Firecracker inside a VM will underperform exactly where agents spend most of their time. Snapshots are bound to the specific Firecracker build that created them, so a fleet running two different builds at once cannot restore each other's snapshots, which is a real concern during any platform upgrade.
The developer-facing layer matters as much as any of the hardware underneath it. Getting from agent code to a live, isolated, persistent endpoint should take a single push, not manual VM provisioning, kernel configuration, or jailer setup handled by hand. That pricing is what the infrastructure underneath it actually costs to run.


