How Firecracker Snapshot-Restore Works for AI Agent VMs
Firecracker trades one slow boot for unlimited fast VM resumptions via frozen state files.

Cold-booting a Firecracker microVM takes time that interactive agent workloads cannot spare: the kernel has to load, init has to run, and the guest agent has to start and signal ready before a single line of agent code executes. Firecracker's snapshot-restore pipeline exists to make that cost disappear from everyday operation by paying it exactly once and resuming from the result in milliseconds. This article walks through how that pipeline actually works, stage by stage, from the files a snapshot contains through the correctness problems that cloning a VM's state creates.
Why Firecracker's Snapshot-Restore Pipeline Exists
Agent workloads do not fit a single mold. Some are short code-interpreter calls that spin up, execute, return a result, and vanish. Others are long-running sessions that hold a persistent workspace, a browser session, or a paused conversation state that needs to resume hours later. What unites them is a shared intolerance for delay at the point of creation: a chat interface cannot absorb a full cold boot every time it spawns a sandbox, no matter how short the resulting task turns out to be. The friction compounds in any agent harness where context builds across many tool calls, because each call may need its own isolated execution environment, and repeating a full boot for each one multiplies a small delay into a real one. Snapshot-restore solves this by separating initialization from execution: getting a VM into a ready state happens once, and the infrastructure resumes that frozen state as many times as the workload demands.
What Firecracker captures in a snapshot and what the files contain
A Firecracker snapshot is a complete record of a running microVM's state at the instant it was taken: the full contents of guest memory, the CPU's register state, and the state of every emulated device attached to the machine. Calling the snapshot API, PUT /snapshot/create, produces two files that hold this record. The first, vmstate.bin, holds CPU registers, device state, and interrupt controller state, and runs to around 16KB regardless of how much memory the guest has. The second, memory.bin, holds the entire contents of guest RAM: a 512MB VM produces a 512MB memory.bin.
The device state inside vmstate.bin reflects Firecracker's deliberately narrow device model. Firecracker supports only six emulated devices: virtio-net, virtio-balloon, virtio-block, virtio-vsock, a serial console, and a minimal keyboard controller used only to stop the microVM. Later versions added virtio-rng and virtio-pmem as optional extras. There is no USB stack, no GPU, no general PCI bus by default, because Firecracker was built for exactly this kind of lean, single-purpose virtualization rather than general-purpose desktop or server emulation. The block device, the rootfs the guest boots from, is tracked separately from memory.bin and needs its own copy-on-write handling at restore time, since every VM restored from the same snapshot needs a disk it can write to without colliding with every other restore sharing that same origin image.
Together, these two files are the entire frozen machine. Nothing about the restore process that follows needs to re-derive state from configuration, because the state itself, down to the CPU registers and the interrupt controller, already sits on disk.
Building the Golden Image: The One-Time Cold Boot That Funds All Fast Restores
Every fast restore downstream depends on one deliberate, slow step taken in advance. The sequence starts with a normal cold boot: a temporary microVM boots from the target rootfs the way any VM would, the kernel loads, init runs, and the guest agent inside starts up and signals that it is ready to accept work. Once that ready signal fires, the operator issues PUT /snapshot/create, which writes vmstate.bin and memory.bin to disk. The temporary VM is then killed, and what remains is the pair of snapshot files plus a clean baseline rootfs, the template everything else will restore from.
That readiness signal is the one correctness requirement in the entire creation step, and it is unforgiving. If the guest agent has not finished initializing when the snapshot is taken, the files capture an unready machine, and every subsequent restore inherits that same unready state. There is no way to patch an incomplete initialization after the fact. The snapshot is a photograph, not a recipe, and whatever was true of the machine at the shutter's click is what every later restore wakes up to.
The economics of this step are what make the whole pipeline worth building. The one-time cost is a single cold boot, and every spawn after that restores in under 30ms. That cold boot is also the last one that image will ever need. Snapshot files are scoped per image: a python:3.12 environment produces its own snapshot pair, a different base image produces a different pair, and an agent harness running multiple distinct environments ends up maintaining a distinct golden image for each one. The snapshot store that results from this is typically organized as one directory per image, keyed on a hash of the rootfs path, holding vmstate.bin, memory.bin, and a clean baseline rootfs.ext4 ready to be copied for each new restore.
How restore works: memory-mapping, relative paths, and the no-config constraint
Restore is fast for a specific mechanical reason: the kernel's own memory-mapping facilities do almost all of the work. There is no kernel boot, no process startup inside the guest, no initialization sequence to wait on. The sequence an operator runs per spawn is short. First, the clean baseline rootfs is sparse-copied to a new directory unique to that sandbox, a copy-on-write operation that completes almost instantly. Second, PUT /snapshot/load is issued with resume_vm set to true, and the VM resumes execution immediately from the exact point it was frozen at. Third, the caller connects to the guest agent over vsock, and because that agent was already running and already signaled ready before the snapshot was taken, it is still running and still ready at the moment of resume. No initialization logic runs a second time.
Getting from one golden snapshot to many concurrent restores depends on a detail easy to miss on a first implementation: the paths baked into the snapshot at creation time. If those paths are absolute, every VM restored from the same snapshot tries to open the same rootfs file and the same vsock socket, and they collide. The fix is to use relative paths when the snapshot is created, for example "path_on_host": "rootfs.ext4" and "uds_path": "v.sock", and to set each Firecracker process's working directory to its own per-sandbox folder before loading. Firecracker resolves those relative paths against the process's working directory, so each restored VM ends up with its own rootfs and its own vsock socket without any extra configuration. Firecracker's own snapshot documentation also allows overriding the vsock path directly at restore through a vsock_override parameter, and drive paths can be patched after load, giving operators a second lever if the relative-path convention alone isn't enough for a given deployment.
A common mistake in first implementations is pausing a VM with a PUT request to /vm with {"state": "Paused"}, when the endpoint requires a PATCH. Sending a PUT where a PATCH is expected fails, and the resulting error gives little indication of what went wrong; developers can lose real time tracing a failure back to a single wrong HTTP verb. Getting the method right the first time avoids a debugging detour that has nothing to do with the actual snapshot logic.
Measured end to end, a single restore breaks down roughly like this: Firecracker process startup takes about 5ms, mapping the memory snapshot file into the process takes about 8ms, restoring CPU and device state takes about 10ms, and vsock reconnection plus the ready signal takes about 5ms, for a total of roughly 28ms. Running three restores from the same snapshot concurrently produced wall-clock times of 28ms, 29ms, and 33ms. Concurrency barely moves the number, because each restore operates on its own process, its own memory mapping, and its own rootfs copy, with no shared mutable state forcing one restore to wait on another.
Copy-on-write memory and disk sharing across restored VMs
Restoring quickly solves latency; restoring cheaply at scale depends on a separate mechanism, memory sharing. When a restored VM maps the snapshot's memory.bin into its address space, it does so with MAP_PRIVATE, the standard copy-on-write mapping mode. That means hundreds of sandboxes restored from the same golden image all read from the same underlying physical pages in RAM for anything they have not touched, and each one only consumes additional physical memory for the specific pages it modifies.
The mechanism runs entirely inside the OS kernel, with no coordination logic required from Firecracker or the operator. Every restored guest reads shared physical pages from the common template. The instant a guest writes to one of those pages, the kernel intercepts the write, allocates a fresh private physical page for that guest alone, and copies the modified content into it. Everything else continues to be shared. Firecracker's own figures put the memory overhead of this scheme at less than 5MB per VM, and the ceiling on how many VMs can run simultaneously from one golden image becomes a question of available hardware, not a limit built into the architecture itself.
The same logic extends to disk. The baseline rootfs.ext4 is sparse-copied, or reflinked on filesystems like btrfs that support it, once per sandbox, and each sandbox's writes land in its own overlay while the base image stays untouched and shared across every restore. Because the snapshot files themselves are read-only once written, concurrent restores need no locking: each restore process opens the same files independently, and nothing about one restore can interfere with another.
Lazy memory loading with userfaultfd: restoring before the full memory image is local
Firecracker extends this model one step further through userfaultfd, a Linux facility that lets a userspace process intercept and handle page faults on a given memory range. Rather than requiring the entire memory.bin to be present locally before a VM can resume, Firecracker can hand guest-memory page faults off to a userspace daemon. The VM can then start executing before its full memory image has finished loading.
The practical effect is that resume_vm can return control to the guest immediately, and individual pages of memory stream in only as the guest actually touches them, whether they live on local NVMe storage, a network-attached volume, or an object store sitting further away. This matters most in deployments where snapshot storage is tiered by how often an agent wakes up: frequently used agents keep their snapshots on fast local disk, while rarely used ones sit in cheaper object storage until needed. Demand-paging through userfaultfd bridges that gap, letting a cold, remotely stored snapshot resume without first forcing a full download to local disk. It is a supporting mechanism rather than the dominant path: most restores in a well-tuned deployment still read a local memory.bin directly, with userfaultfd reserved for the distributed and storage-tiered cases where it earns its complexity.
The uniqueness problem: what cloning a snapshot does to cryptographic state, UUIDs, and nonces
Restoring the same snapshot many times means restoring the same memory contents many times, and that creates a hazard with no mechanical fix at the restore layer itself. Anything in guest memory that is supposed to be unique, cryptographic secrets, random number generator seeds, UUIDs, nonces, is identical across every clone the instant they are restored, and the guest has no built-in way of knowing it has been cloned.
The resulting failures do not look like crashes. Two VMs restored from the same snapshot can generate the same UUID, because both inherit the exact same PRNG state that existed at the moment the snapshot was taken, and the result is a silent correctness failure with no error message. Cryptographic key material sitting in RAM when the snapshot was created sits in every clone's RAM when each one restores, so an attacker able to read the memory of one clone can infer the secrets of every other clone descended from the same snapshot. TLS session state, CSRF tokens, and any value derived from random bytes generated before the snapshot point all get duplicated the same way.
Research from Amazon Web Services, published by Marc Brooker, Adrian Costin Catangiu, Mike Danilov, Alexander Graf, Colm MacCarthaigh, and Andrei Sandu, addresses this directly in the context of serverless cold-start, where post-initialization memory snapshots are cloned and restored constantly and the uniqueness problem is unavoidable at scale. The paper proposes two interfaces built to close the gap. MADV_WIPEONSUSPEND lets the guest OS mark memory ranges holding high-value secrets before suspension, so the kernel zeroes those specific pages as part of the snapshot, and every clone starts without the original secret material, forced to derive fresh secrets on resume. VmGenId takes a different approach: it exposes a 128-bit generation identifier to the guest as a virtual device, and the specification states that this value is a cryptographically random identifier that changes whenever a virtual machine is cloned or restored to an earlier state, a fresh random replacement on every restore. A guest written to watch for that change can detect that it has been cloned and re-seed its random number generator and any other uniqueness-sensitive state accordingly.
Neither mechanism is automatic. Both require the guest operating system and application code to actually use them, mark the right memory ranges, and watch for the generation identifier's change. The tools exist and function as designed; the risk is assuming a restore is safe by default when no one has wired those tools into the guest.
Version coupling and CPU template compatibility: the production risks that don't announce themselves
A snapshot is a serialized record of one specific machine configuration, down to the Firecracker version that produced it and the exact instruction set of the physical CPU the original VM ran on, and restoring it elsewhere can fail in two very different ways. The first is a loud failure: the restore call simply errors out because the new VMM version cannot parse the structure of an older vmstate.bin, and the operator finds out immediately. The second is worse, because it is quiet at first. A snapshot restored onto a physical CPU with a slightly different instruction set can load successfully, resume successfully, and run correctly for minutes before the guest executes an instruction the new host doesn't actually support, at which point the VM falls over with no warning tied back to the real cause.
This class of failure is operationally more dangerous than the quiet divergence of identical images across hardware, precisely because it hides. A missing re-seed of a random number generator is a correctness bug that a careful audit can find by reasoning about what the guest does with randomness. A CPU instruction mismatch is a landmine that depends on which specific instruction the guest happens to execute and when, so two deployments running the ostensibly same image on different hardware generations can diverge in behavior with no code change on either side. Matching CPU templates between the host that created a snapshot and the host that restores it, and pinning the Firecracker version across both ends of the pipeline, is the condition under which the rest of this pipeline's speed and efficiency claims hold.


