Est.

Idle Agent Cost at Scale for SaaS Products

Idle compute between model calls drives runaway agent costs, not the models themselves.

Senior Writer · · 10 min read
Cover illustration for “Idle Agent Cost at Scale for SaaS Products”
Agent Hosting Economics · October 1, 2026 · 10 min read · 2,285 words

The real driver of runaway AI spend in production is the compute that sits idle between model calls. It's the compute that sits idle between those calls, billed in full while the agent waits for a tool to respond, a user to confirm, or a retry to resolve. Uber's engineering organization offers the clearest record of what happens when this gap goes unmanaged. Claude Code adoption swept through nearly the entire engineering organization between December 2025 and March 2026, and by April the entire annual AI budget was gone. Typical monthly API costs per engineer ran into the hundreds of dollars, while heavy users reached $500–$2,000. The shock didn't come from the pilot, where usage was narrow and costs looked manageable. It came from production, where agents ran constantly, waited constantly, and billed for both.

The compounding structure behind agent costs outgrowing user counts

Diagram: Why Agent Costs Compound Instead of Scale. Visualizes: Illustrate the difference between how costs grow with users for a standard chatbot versus an agentic workflow.

Uber's budget collapse was a symptom of a structural problem, not a one-off accounting failure. It reflects a structural property of how agents work: a single user request doesn't trigger one model call, it triggers a chain, planning, retrieval, tool calls, validation, retries, and synthesis, and each link in that chain carries its own context package and its own cost. Adding users multiplies cost rather than adding it in a straight line. You multiply it, because every new user brings a new chain, and every chain can branch into retries that look nothing alike even when the underlying task is identical.

That unpredictability is a property of how agents reason and does not average out at scale. It's a property of how agents reason: ambiguity in a request forces more planning steps, a failed tool call forces a retry, and a retry forces the whole context package to be rebuilt and re-sent.

It would be reasonable to assume that falling model prices offset this, but they don't. Enterprise AI inference now makes up the dominant share of total AI budgets, and agentic workflows consume many times more tokens per task than a standard chatbot query. Token consumption is growing faster than unit prices are falling, so the net effect on the bill is upward. The gap between a demo and a production deployment is where this becomes visible in dollar terms. A support, finance, or data agent that looks inexpensive against a narrow, scripted demo meets a different reality once it has to handle ambiguous requests, missing context, and permission checks at real volume, and that is where the total cost of ownership actually breaks, not in the sticker price of the model.

What persistent state requires (and why stateless infrastructure makes idle costs worse)

Large language models have no memory of their own. Every agent built on top of one has to reconstruct context on every call, and whether that reconstruction turns idle time into billable waste depends on the infrastructure decision behind it.

Five architectural patterns exist to give an agent something resembling memory: in-context working buffers that use sliding windows with summarization, execution checkpointing to durable storage, semantic memory held in vector stores, episodic event logs, and multi-scope memory isolation for multi-tenant deployments. All five share one requirement: the agent's environment has to survive between sessions for any of them to work. Frameworks such as LangGraph have built persistence into their core design, with typed state, conditional edges, checkpoint-based resumability, and interrupt points that let a human step in before a risky action executes. None of that functions unless the infrastructure underneath is actually holding the agent's state while it isn't running.

Gating every tool call behind a persistence checkpoint lets a workflow resume exactly where it stopped instead of replaying everything that came before. Gates can be scoped narrowly, so an agent runs freely through low-risk steps and only pauses where a human needs to approve something consequential. That design only works if the execution environment is persistent rather than a function that spins up, runs, and disappears.

This is where the obvious fix, go stateless, go serverless, falls apart on inspection. A stateless-serverless model does eliminate idle billing, but only by eliminating state itself, and the cost doesn't disappear, it moves. It reappears as cold boot time, as the work of reconstructing context from scratch, and as context that has to be re-sent on every resume. Those are exactly the costs driving the compounding bill described above: more tokens, more steps, more variance. Eliminating idle billing by eliminating state solves the wrong problem. The real requirement is an architecture that keeps state intact while still refusing to bill for the time an agent spends doing nothing.

How snapshot-restore turns isolation into an economics primitive

Diagram: Snapshot-Restore: Compute Billed Only When Active. Visualizes: Show the lifecycle of a Firecracker microVM agent under snapshot-restore scheduling: boot once (~125 ms), run active work (billed as compute), snapshot to storage (5–30 ms)…

Snapshot-restore is the mechanism that resolves this tension. It takes micro-VM isolation, originally a security feature, and turns it into an economics primitive: pause an agent, write its memory and filesystem state to storage, and restore the whole thing in milliseconds. While the agent is dormant, it costs the price of storage, not the price of compute.

Firecracker microVMs are the technology that makes this practical. They boot in roughly 125 milliseconds, carry very low per-VM memory overhead, and support snapshot-restore that captures both memory state and block device state together. That snapshot becomes, in effect, the agent's identity across sessions, the thing that gets restored rather than rebuilt. The restore path is fast enough that users never notice it happening: Firecracker's snapshot-restore completes in the 5 to 30 millisecond range, so a dormant agent is awake and responding before a human could perceive any delay. The billing model changes completely. The product experience does not.

The isolation boundary is not incidental to this design; it's what makes running arbitrary agent code safe. A standard container shares the host kernel, so a single agent that escapes its sandbox can compromise everything else running on that host. A microVM gives each agent its own guest kernel behind a hardware virtualization boundary, the same approach AWS Lambda uses in production. That boundary is what lets an agent install packages, open a browser session, or run code it wasn't explicitly told to run, without putting any other tenant's agent at risk.

What follows from this is a scheduling model: boot a microVM to a ready state once, snapshot it, and restore every subsequent session from that snapshot instead of cold-booting from zero. Infrastructure cost depends on how often a VM is actually doing something, not on how long it stays up. Snapshot-restore does not make infrastructure free. It removes the charge for the hours an agent spends waiting, which, as the Uber numbers show, is often most of them.

Idle Billing Costs at Scale

Wall-clock billing charges for time elapsed, not work done. An agent that executes code for a few seconds and then waits an extended stretch for a user's confirmation pays for the full span, whether it's computing or sitting idle, because the sandbox is metered by the clock rather than by activity.

An August 2026 benchmark that priced realistic agent workloads found a meaningful cost spread on identical work, and the spread tracked idle policy more than anything else about the workload itself. The cheapest outcome in that benchmark came from active-CPU metering rather than wall-clock billing, showing that the architecture governing how idle time is charged produced the gap. At production scale, the calculus shifts again: a cost comparison found roughly an 8× delta between managed platform-as-a-service billing and bring-your-own-cloud arrangements for the same workload, a gap attributable almost entirely to the margin embedded in managed per-second pricing.

The market has started building around this directly. Several platforms now compete explicitly on eliminating idle charges rather than on raw compute price. Blaxel advertises sandboxes that remain in standby indefinitely with sub-25ms resume and zero compute cost while idle. Comparable patterns, pause-and-fork on one platform, checkpoint-and-restore on another, point to the same underlying design principle: treat the snapshot as the unit of state, and agents get branching, resumable workflows instead of repeated cold boots.

The lesson for anyone building a SaaS product on agents is that idle cost isn't something to tune after the product ships. It gets set the moment an architecture is chosen, before the first user ever logs in. A product built on wall-clock billing can optimize prompts, cache context, and tighten retries, and it will still carry a structurally higher idle cost than one built on activity-based billing from the start.

Economic viability of per-user agent deployment at near-zero idle cost

Agent-native SaaS products are converging on a particular shape: one persistent agent per customer, each with its own memory, its own credentials, and its own execution environment. This is the natural unit for products that need to remember a specific user's history, preferences, and prior decisions rather than treating every request as a cold start. The problem is that most of those agents spend most of their time doing nothing. A fleet of thousands of per-user agents, each idle the overwhelming majority of the day, is a fleet of thousands of wall-clock bills if the underlying infrastructure doesn't distinguish idle from active.

Snapshot-restore scheduling is what makes that fleet affordable. A dormant agent pays for storage rather than compute, so the per-user cost converges toward the cost of holding a disk snapshot. A fleet of thousands of user-specific agents stays viable this way, while a fleet of always-on VMs becomes a budget line that grows without bound.

Pricing at the SaaS layer is already moving in a direction that assumes this is possible. Seat-based pricing fell meaningfully as a share of SaaS companies between 2025 and 2026, while hybrid pricing models surged over the same period. Charging per seat means charging for headcount regardless of whether those seats generate any value that day. Agent products need pricing tied to actual activity, because the underlying cost structure, done correctly, is also tied to actual activity. Outcome-based pricing at the product layer holds together only if the infrastructure underneath bills for what ran, not for what merely existed.

Building a per-user agent fleet: the architectural decisions that determine whether the economics work

Every agent in a fleet needs auth isolation, memory isolation, and cost isolation at the same time, and these constraints have to be solved together, not in sequence. Solve only one and the fleet ends up insecure, amnesiac, or economically unworkable.

On memory, a few choices recur across serious implementations. Execution checkpointing to durable storage, a system like PostgreSQL, for example, gives fault tolerance and resumability, so an agent process can be killed outright and restored without losing the progress of whatever workflow it was running. Semantic memory in a vector store supports recall across sessions, and episodic event logs let an agent learn from its own past failures, but both require a storage layer that persists independently of whatever compute happens to be running at a given moment. Each user's agent needs to read only its own memory, which maps to per-VM storage isolation rather than to a shared database protected by access controls layered on top.

Naive adoption of MCP tooling carries a real cost that's easy to miss in a demo: connecting a standard GitHub MCP server can consume approximately 55,000 tokens in context before the agent does anything useful, which makes tool-definition overhead a genuine line item in total cost once a fleet is running at scale. The production-grade pattern for multi-tenant fleets is white-label per-user auth, where each customer authorizes their own account inside the product's own branding and the resulting tokens stay server-side rather than passing through the client. MCP's OAuth flow is a standard with an authorization model that survives more than one user, carrying per-user auth so the server acts as whoever is asking and can log it.

On infrastructure, snapshot-restore scheduling belongs in the initial design, not on a list of things to optimize once the bill arrives. Idle agents should pay for storage, not compute, by default, which makes the fleet's cost floor a function of storage volume rather than active VM count. Wake latency under one second is a hard requirement for a responsive agent product, and a restore path in the 5 to 30 millisecond range clears that bar comfortably enough that users never perceive the gap between dormant and active. Isolation at the kernel level, each agent running in its own microVM, is what lets agents install packages, open browser sessions, and run arbitrary code without putting the rest of the fleet at risk, and that property is what separates a real execution environment from a sandboxed function that merely looks like one.

Economic viability of snapshot-restore scheduling for fleet deployment

Snapshot-restore scheduling delivers its biggest gains where agents spend the overwhelming majority of their time idle. A user-facing agent that runs one session a day and sits dormant the rest of the time pays almost nothing in compute under this model, and the economics get stronger, not weaker, as fleet size grows and per-agent activity stays low. This is the profile most per-user SaaS agents actually fit: bursts of real work separated by long stretches of waiting for the next request.

The model offers less advantage for agents that are built to run continuously. A pipeline agent processing a batch job for hours without interruption isn't paying an idle tax in the first place, so there's little idle cost for snapshot-restore to eliminate, and the overhead of pausing and restoring a workload that was never going to pause anyway buys nothing. The right question for any team evaluating this architecture is what fraction of the fleet's total running time is actually idle, because that fraction shows whether the architecture is the right fix or an answer to a problem the workload doesn't have.

Sources

  1. Cost to Run AI Agents at Scale: The Full TCO Breakdown 2026
  2. The Bill Arrives: How to Manage Agentic AI Costs at Scale