Sub-Second Agent Wake Latency Requirements and Benchmarks
Millisecond wake delays determine whether persistent agents feel responsive to users.

Wake latency is the delay a persistent, stateful agent incurs before it can do anything useful after a period of dormancy, and it has no real equivalent in the stateless request-response systems most engineering teams already know how to operate. A stateless API pays its startup cost once, at deploy time. Every call after that hits a process that is already warm, with its memory loaded, its connections open, and its runtime ready to execute. An agent that sleeps between sessions does not get that luxury. It must re-establish its entire execution context, memory state, open connections, loaded tools, on every single wake, because the whole premise of a persistent agent is that it goes idle between interactions and has to come back to life each time a user returns. This distinction matters because it changes where the cost falls in the system. In a stateless architecture, the cost is front-loaded and amortized across the service's lifetime. In a persistent agent architecture, the cost recurs with a frequency tied to how often users interact with the agent, not how often the infrastructure team redeploys it. That recurrence is the structural fact the rest of this piece builds from.
How human perception thresholds turn milliseconds into product failures
The ceiling on acceptable wake latency is not a matter of taste or product polish. It comes from how human conversational turn-taking actually works, and the thresholds are tight. Pauses that approach or exceed roughly one second already read as unnatural in a live exchange, registering as hesitation. Pauses that stretch past two seconds mark a clearer communicative breakdown, the kind that signals to a human participant that something has gone wrong. These thresholds hold regardless of the interface. They apply to voice agents taking a caller's next instruction, to chat agents composing a reply, and to any interactive loop where a person is waiting on the other end. None of this is specific to a particular platform or vendor's implementation, because it describes something about human perception rather than something about software.
The consequence for agent infrastructure is immediate. If an agent's wake cost alone consumes several hundred milliseconds, that cost is incurred before the system has generated a single token of actual reasoning or response. The agent starts the race already behind the one-second mark, and it has not yet done anything a user would recognize as useful work. Inference time, tool calls, and reasoning steps all still have to happen after that. A wake penalty of a few hundred milliseconds is a meaningful fraction of the entire one-second perceptual budget, spent before the agent has begun to think.
How multi-step reasoning loops make wake latency compound
Production agents rarely answer in a single inference call. They run a trajectory, a sequence of reasoning steps, each of which consumes real time on its own. A single reasoning cycle in a production agent already costs somewhere between one and several seconds of pure LLM inference. A task that requires multiple reasoning steps stacks those cycles sequentially, so the total inference budget for a moderately complex task can reach several seconds before any wake cost is even added to the ledger. A few hundred milliseconds of wake cost does not just tack onto the end of a multi-second trajectory as a flat tax. It is paid at the front, before the first step runs, and it then colors the user's perception of every step that follows, because the user experienced the delay as the opening beat of an interaction that was already going to take several seconds to complete.
This dynamic gets sharper once reflection and planning patterns enter the picture. Agents that critique and revise their own outputs, running an extra reasoning round specifically to improve quality, are a known and often worthwhile trade of latency for better results. But that trade only makes economic sense if the baseline wake cost going into the trajectory has already been minimized. Adding deliberate extra reasoning rounds on top of an unoptimized wake penalty stacks two avoidable costs together.
The structural reason wake cost sits inside the trajectory rather than outside it comes down to how tool calls actually execute. Each tool call an agent makes requires launching an isolated execution environment, waiting for that environment to produce a result, and feeding the result back into the next reasoning step. Sandbox scheduling latency, the time spent waiting for that environment to become ready, is not background overhead running in parallel with something else. It sits directly on the critical path of the reasoning loop itself. Minimizing it has to be treated as a first-order architectural concern.
How fleet economics convert latency into an infrastructure cost
An agent that cannot resume cheaply has exactly two bad options. It can cold-boot on every invocation, which is slow and violates the perceptual thresholds described above. It can stay resident in memory indefinitely to avoid that cold-boot penalty. That means paying for compute the agent is not using during every idle moment between user interactions. Neither option scales, and the reason becomes obvious the moment the system has to serve more than a handful of agents at once.
A production fleet has to support many concurrent agent sessions, and the overwhelming majority of those agents are idle at any given moment, waiting for their user to send the next message or take the next action. Keeping every one of them warm in memory to dodge startup latency means holding compute allocation for the entire fleet at all times, even though only a small fraction of it is doing anything productive at any given moment. That cost scales directly with fleet size, and fleet size is set by something the infrastructure team does not control: how many users the product has.
The per-user agent model, where every user gets a dedicated, persistent agent rather than sharing a pool of generic workers, makes this tension unavoidable. Fleet size scales with the user base, and idle compute cost scales right alongside it. Whether to cold-boot or stay warm becomes the defining economic question of any multi-tenant agent deployment, and it has no good answer within the constraints of either option alone. The only way out of that trade-off is an architecture that makes agents cheap to suspend when idle and fast to resume when needed, so that the fleet is not paying for idle compute and is not paying in latency either. Snapshot-restore VM scheduling is built to deliver exactly that combination, and it is worth understanding in some mechanical detail why it can.
Cold-boot VM latency and the numbers behind it
Cold-boot latency sets the baseline that any fix has to beat, and even well-engineered virtualization falls short of that bar. Firecracker, the microVM technology purpose-built for fast boot and widely used precisely because of its minimalist design, still runs through a cold-boot sequence with several irreducible phases: loading the kernel, running init, and starting the agent's own runtime. Each of those phases takes real time, and together they add up to a total that exceeds what an interactive agent turn can tolerate when that cost has to be paid from scratch.
A boot time of around one second is entirely acceptable for a long-running workload that starts once and then runs for hours or days. It is a different matter when that same cost has to be paid on every single agent invocation inside an active session, because a user is not waiting once for the system to come online, they are waiting every time they send a message or trigger a new tool call. Sandbox launch latency, the time needed to get an isolated execution environment ready for the agent's first tool call, is a real and significant source of overhead in agent serving. The deeper issue, though, is architectural rather than purely mechanical: much of the end-to-end delay in agent serving traces back to the sequential execution model used by conventional agent runtimes, where each step waits for the previous one to fully resolve before starting. Fixing sandbox initialization alone does not fix that sequential bottleneck, which is one reason the solution to wake latency has to operate at the level of the execution model, not just at the level of individual VM boot times.
Snapshot-restore: from sequential startup to instant memory resume
Snapshot-restore does not make the cold-boot sequence faster. It removes the need to run that sequence. Instead of loading a kernel, running init, and starting the agent's runtime in order, a restore operation memory-maps a previously saved VM state and resumes execution from the exact instruction where the agent was suspended. Every sequential phase that cold boot has to walk through gets bypassed entirely: the work those phases do was already done once, at snapshot time, and the result was saved.
From the perspective of the guest system itself, the VM was never off. Its memory contents, its open file descriptors, its running processes, and its network connections are all exactly as they were at the moment of suspension. The only thing that changed is that wall-clock time advanced while the VM sat dormant. That property is what lets snapshot-restore produce wake latencies measured in tens to low hundreds of milliseconds at the hypervisor level, with full end-to-end wake latency typically landing in the low-to-mid hundreds of milliseconds overall, comfortably inside the perceptual thresholds described earlier.
The benefit compounds because what gets restored is not just a generic VM shell but the agent's actual working state: installed packages, active credentials, open browser sessions, and whatever context the agent was holding in memory before it went idle. None of that has to be reconstructed or re-initialized on wake. The agent resumes doing useful work immediately, with zero re-initialization logic standing between the restore completing and the next reasoning step executing. Fast sandbox scheduling, including techniques like speculative prewarming and prefetching that prepare environments before they are requested, keeps that restored state ready to go instead of sitting behind a queue, and that scheduling layer is what ultimately keeps wake latency off the reasoning loop's critical path.
Pre-allocated networking and pool-based scheduling under concurrency
Snapshot-restore solves the problem of VM startup time, but it does not automatically solve what happens to networking when concurrency rises. A restore operation that completes in tens of milliseconds for a single VM can stretch into seconds once hundreds of VMs are restoring at the same time, if each one has to set up its own networking from scratch. The cause is straightforward: creating network namespaces, configuring interfaces, and applying routing rules is real work, and that work does not parallelize cleanly. At high concurrency it serializes into a queue, and the resulting delay grows in proportion to how many VMs are starting at once.
The fix follows the same logic snapshot-restore already applies to VM state. Do the expensive setup work ahead of time, hold the finished result in a pool, and let restore connect a VM to an already-configured network slot. This turns networking from a per-restore cost into a one-time setup cost amortized across the pool, the same trick that made snapshot-restore itself viable for VM state. Speculative pre-warming, predicting which agents are likely to be invoked next and restoring their snapshots before the request even arrives, is what converts this from a latency improvement into something closer to a zero-latency primitive at fleet scale. A sub-second wake latency guarantee is achievable at scale only when both pieces are in place at once: the scheduler pre-warming snapshots ahead of demand, and the network stack pre-allocated and waiting. The VM technology makes sub-second wake possible. The scheduling layer is what keeps it true under real production load.
Micro-VM isolation as the security model for agents executing arbitrary code
Agents that run scripts, call tools, install packages, and operate browsers are executing arbitrary code on behalf of a user, and that capability demands a security boundary strong enough to contain a compromised agent before it can reach the host system or another agent's environment. Hardware-virtualized micro-VMs provide that boundary in a way container-based and serverless runtimes cannot, because each micro-VM runs its own kernel on top of hardware virtualization such as KVM. A kernel exploit inside one VM has no path to the host or to a neighboring VM. Containers, by contrast, share the host kernel, so a kernel-level compromise in one container can potentially reach every other container on that host.
The stakes of that isolation boundary rise sharply once the thing being isolated is an autonomous agent. An agent executing a multi-step plan with access to tools may be holding live credentials, open sessions, and accumulated context from earlier in its trajectory, all of which becomes available to an attacker the moment that agent's execution environment is breached. Hardware-rooted runtime policy enforcement frameworks built specifically for agentic systems exist because the blast radius of a compromised agent is categorically larger than the blast radius of a compromised stateless API call.
The per-user agent model makes this an unavoidable product requirement. If each user's agent holds that user's credentials, files, and session history, then any mixing of two agents' execution environments, however brief or accidental, constitutes a data breach between two customers. Micro-VM isolation is the architecture that makes running a dedicated, persistent agent per user safe by default.
Speculative execution and action-observation co-speculation below the reasoning loop's critical path
Snapshot-restore and pre-allocated networking establish a floor for wake latency, but the most advanced agent serving systems do not treat that floor as the end of the optimization. Instead of waiting for a reasoning step to finish before starting to prepare the next execution environment, these systems speculate on what the agent is likely to do next and begin preparing that environment while the current step is still running. This hides wake latency behind inference time.
Action-observation co-speculation, the technique introduced under the name AOSpec, takes this a step further by speculating on both sides of the loop at once. The serving system predicts the agent's likely next action and executes it speculatively inside an isolated fork, while simultaneously drafting the likely observation or tool output the environment would produce, using a technique called Expected Value Decoding. Both predictions run while the current inference step is still in progress. If the predictions turn out correct, the sandbox and its expected output are ready the moment the agent actually asks for them. If the predictions are wrong, the speculative work is simply discarded and the system falls back to the normal path. What this does structurally is convert wake latency from a sequential cost, paid only after each reasoning step completes, into a parallel cost, paid during the step that precedes it. On correct predictions, the agent experiences something close to zero sandbox latency, because the waiting happened earlier, hidden inside computation that was going to happen anyway.


