Est.

Storage vs Compute Billing for Dormant AI Agents

Snapshot-restore cuts idle compute waste by trading it for cheaper dormant storage.

Contributing Editor · · 11 min read · Updated
Cover illustration for “Storage vs Compute Billing for Dormant AI Agents”
Agent Hosting Economics · October 2, 2026 · 11 min read · 2,501 words

Cloud compute billing was built for workloads that stay busy, and agent workloads don't stay busy: that mismatch is why so many teams open an invoice and find a number several times larger than what they planned for. In April 2026, GitHub paused new Copilot signups after a handful of agentic requests burned through what users had paid for entire monthly subscriptions. That wasn't a pricing mistake on GitHub's part. It was the visible edge of a structural problem that every agent platform runs into eventually.

Most cloud infrastructure still bills the way it did for web servers, which process requests continuously across long sessions that stay open for minutes or hours. Agents don't behave that way. Blaxel's 2026 analysis describes the agent cycle as: spin up, run a task that takes seconds, shut down, then repeat the cycle many times within a single user session. The meter, however, doesn't know the difference. Between invocations the sandbox sits idle, but a traditional VM or container keeps charging for the full allocation window it was given, whether or not anything inside it is doing work.

Three separate cost problems stack on top of each other as a result. The sandbox burns money while it waits for the next instruction. Billing floors punish short tasks, so a job that finishes in two seconds can cost the same as one that runs for the full minimum increment. And teams over-provision for peak load, because agent memory use can spike to several multiples of its baseline during a burst of tool calls, so capacity planning gets built around the worst moment instead of the average one. Average Kubernetes CPU utilization already runs around 8%, by Blaxel's accounting, and agent sparsity pushes that number down further still. Capacity sits provisioned and idle while the workload underneath it does almost nothing most of the time. That's the gap between what teams estimate a system will cost and what the invoice says.

Why teams accept the waste: the cold-start trap

The obvious response to idle compute is to shut the sandbox down between invocations. Teams mostly don't, and the reason is rational: cold starts wreck the experience for anything interactive. A real-time coding assistant, a PR review agent, a sequential tool-call workflow, these are exactly the highest-value agent products reaching production right now, and they're also the ones users notice the lag on first. Nielsen's long-standing research on response time, which Blaxel cites directly, puts the ceiling for something to feel instantaneous at 100 milliseconds. A multi-second cold boot is an order of magnitude past that number, in the range where users register the delay as the system being broken rather than the system being slow.

So teams keep sandboxes warm. They eat the idle compute bill because the alternative, a visible stall every time a user re-engages an agent, costs more in churn and frustration than the wasted infrastructure spend costs in dollars. That decision makes sense given the choice as it's usually presented: pay for idle capacity, or make users wait. Framed that way, keeping the sandbox hot is the only defensible option.

The trap is that the choice itself is artificial. Blaxel frames it as a kind of lock-in, where the billing model forces a pick between performance and cost that better architecture simply removes. Shared infrastructure forces a full cold boot to isolate one agent's state from another, manufacturing this binary. Each agent running in its own persistent micro-VM with snapshot capability, as platforms like Maritime demonstrate, dissolves the constraint, because wake latency under a second becomes possible once the agent doesn't need to spin up a new container or kernel. It only has to restore a frozen state that was already sitting there, waiting.

Snapshot-Restore: Mechanism and Cost

Snapshot-restore turns the cold-start-versus-idle-waste binary into a three-way choice, but you need to know what it actually does before you weigh it against the alternatives. Pausing a sandbox into standby preserves the full filesystem and memory state, including whatever process context was running at the moment of the pause. Resuming restores that exact snapshot. Compute charges for CPU and memory drop to zero for the entire standby period, because nothing is running, nothing is being provisioned, and nothing is billable in the traditional sense.

A detail in how memory gets mapped on restore is what makes this fast enough to use in production. Snapshot memory is mapped MAP_PRIVATE: the guest shares the snapshot's pages read-only until it actually writes to one of them, at which point only that single page gets copied rather than the whole memory image. That's why restoring a snapshot doesn't mean duplicating gigabytes of RAM every time an agent wakes up, and it's the reason sub-second wake times are achievable instead of theoretical.

Production numbers back this up. PandaStack reports restore times around 49 milliseconds, with end-to-end create times at p50 of 179 milliseconds and p99 around 203 milliseconds. A true cold boot only happens once, on the first spawn of a given template, at roughly 3 seconds, and every create after that pays restore prices instead of cold-boot prices. Research on an idealized version of this scheduling approach suggests it can cut total cost substantially relative to a baseline of keeping sandboxes persistently running, and the distance between what's achievable today and that ideal is largely a scheduling problem rather than a hardware one.

What snapshot-restore does not do is make the cost disappear. It converts one kind of charge into another. Idle compute becomes idle storage, and whether that trade is a good one depends on how storage behaves as a cost, which is a separate question from whether the mechanism works. The per-agent Firecracker micro-VM model keeps each agent's snapshot isolated and independently copyable, with no entanglement between one agent's frozen state and its neighbor's. That isolation is what makes the storage-versus-compute tradeoff something you can actually see and reason about at the infrastructure layer, rather than something buried inside a shared pool.

Why storage billing is structurally cheaper for dormant agents

Compute billing charges for provisioned capacity whether the agent does anything with it or not. Storage billing charges for what the snapshot occupies on disk, a quantity that's bounded by the size of the agent's state and doesn't grow just because the agent sits idle longer. That distinction is the center of the whole argument: paying for compute during dormancy is paying for waste, while paying for storage during dormancy is paying for something that actually exists and holds steady.

Google's own documentation on spend caps makes the asymmetry explicit. A spend cap blocks new usage, but as Google states it, "ongoing fixed usage tied to persistent resources such as compute and storage remain active and continue to accrue charges." Compute during dormancy is waste that varies with how long a sandbox happens to sit around doing nothing. Storage during dormancy is fixed, in the plain sense that the bill doesn't move just because nothing is happening.

Storage costs are climbing, and that argument deserves a real objection. Lucidity's analysis documents a wave of block storage changes from AWS, Azure, and Google across the first half of 2026, all shaped by AI workload demand, adding new storage tiers and performance ceilings that push average storage spend upward. Lucidity also reports that average enterprise block storage utilization is between 15 and 30 percent, so most of what enterprises already pay for in storage sits unused, and AI workloads are making that worse, not better.

The objection holds, but it describes a different failure than the one snapshot-restore billing creates. Lucidity's waste comes from over-provisioned volumes: capacity gets bought for a peak that rarely arrives, and it sits there regardless. A snapshot-based model pays for what you actually write to disk, not for a volume sized against a hypothetical maximum. Enterprise block storage waste is a provisioning discipline problem. Dormant-agent snapshot storage is a different, narrower, and genuinely bounded cost.

A second objection matters just as much, and it's about security rather than price. Snapshot-restore isn't neutral in what it does to an agent's history. Research by Zheng et al. (2026) describes semantic rollback attacks in agent checkpoint-restore systems, where a naive restore can duplicate effects that were already consumed, or bring back authorities and permissions that had since lapsed. The fix is to record which effects an agent has taken that can't be undone, and to choose deliberately between replaying a session or forking a new one from the snapshot, depending on what state actually needs to persist. Storage billing only works as an economic argument if the infrastructure underneath it takes that distinction seriously, a scheduling and systems-design problem as much as a pricing one.

What the infrastructure underneath storage billing requires

Storage billing isn't a pricing decision a platform can simply announce. It depends on specific properties in the infrastructure that most general-purpose cloud compute doesn't provide by default, and some architectures can carry the model while others can't.

Serverless functions can't support it. They're stateless by design, with no persistent disk to speak of, and their cold starts are exactly the problem a storage-first model is meant to solve. Paying for storage makes no sense when there's nothing persistent sitting on disk to store.

Shared containers run into a different wall. Tenants on a shared container share kernel space, so one agent's snapshot ends up entangled with state belonging to other workloads running alongside it. Isolation has to come first, because snapshot semantics only mean something when a snapshot is a complete, self-contained unit rather than a partial slice of a shared system.

The micro-VM model solves this by giving each agent its own kernel, its own filesystem, and its own memory space, so a snapshot is fully self-contained and can be restored without coordinating with anything running nearby. AWS Bedrock AgentCore is the most complete public reference for what this looks like in practice: each session runs in an isolated microVM with its own filesystem access, interactive shell access was added as a further capability in June 2026, execution windows run up to 8 hours compared to Lambda's 15-minute cap, and MCP servers maintain session context across interactions as of March 2026. Google's ADK aims at a related goal in a different way: an event-driven Agent Runtime built around session persistence, auto-scaling that includes scale-to-zero during idle periods, and Cloud Trace integration for observability.

Maritime was built from the start around the architecture storage billing depends on, not retrofitted onto a general-purpose platform. Each agent runs in its own Firecracker micro-VM on bare metal, with a persistent disk, its own kernel, the ability to install packages, its own browser sessions, and wake-from-sleep timing under a second. The snapshot-and-restore scheduler charges for storage during dormancy instead of compute, which is what makes deploying one agent per user economically workable at scale rather than a cost center that grows without bound.

How per-user agent fleets change the economics at scale

At the scale of one agent, idle compute billing is an annoyance, a few extra dollars on a monthly invoice that nobody bothers to chase down. At the scale of thousands of per-user agents, most of them dormant most of the time, that same annoyance becomes the single largest line item on the bill. The per-user agent model, where every user gets a dedicated agent rather than sharing a pool, multiplies the cost of getting this wrong just as much as it multiplies the benefit of getting it right.

Multi-tenant deployments can cut infrastructure costs substantially compared to giving every tenant a fully siloed deployment, but you pay for that savings with a tradeoff in isolation complexity. The practical floor for SaaS multi-tenant deployments is micro-VM or gVisor-level isolation, paired with separate vector namespaces per tenant and per-tenant credential vaults, so that cost savings from sharing infrastructure don't come at the expense of one tenant's data leaking into another's.

A related problem occurs once multiple sub-agents share a single tenant's quota. Several sub-agents run concurrently and compete for the same pool of capacity, so if there's no coordination, each one can stay inside its own rate limit while the group together exhausts the tenant's shared allowance. The fix is a shared ledger per tenant, where each agent registers a soft reservation against the pool and idle agents donate their unused capacity back for others to use.

Session-level spend controls, introduced in August 2026, address a related but distinct problem. Anthropic shipped a hard dollar cap on individual Claude Managed Agents sessions, and AWS added temporal policies to the AgentCore gateway, and both aim to stop a single runaway session from spending without limit. NerdLevelTech's analysis of both announcements notes that Anthropic enforces its cap between model requests, so a session can still overshoot its cap by the cost of one final request before the cutoff takes effect. These caps are a governance layer sitting on top of compute billing. They bound how much damage one stuck agent can do in a single session, but they do nothing to remove the idle compute charges that accumulate between sessions when the agent isn't stuck at all, just waiting. Maritime's flat-rate model works from a different premise entirely, that most agents are idle most of the time, and that charging for storage during that dormancy makes the per-user agent model viable in a way per-minute compute billing can't match, no matter how tightly the spend caps are set.

What builders should verify before committing to a billing model

A platform's billing page only means what it says if the infrastructure underneath it can actually pause, isolate, persist, and restore an agent at the speed and granularity the pricing model assumes. If you're evaluating a platform, or auditing your own stack, you should test the mechanism directly instead of taking the pricing claim at face value.

Wake latency should be tested under realistic concurrent load. A sub-second restore time measured for a single agent waking up alone doesn't tell you whether that number holds when a thousand agents wake at once. If restore times degrade under concurrency, the storage-billing economics fail exactly when they need to hold up most.

Snapshot isolation needs a direct answer, not an assumption. If a snapshot shares kernel state with neighboring agents, it isn't a complete, trustworthy unit of state, and restore behavior under those conditions becomes unpredictable in ways that are hard to debug after the fact. The specific question to ask any platform is whether agents run in separate micro-VMs with their own kernels, or whether they share kernel space with other tenants underneath an abstraction layer.

Finally, "idle" needs a precise definition before anyone signs a contract around it. Some platforms charge a minimum fee per invocation no matter how short the task runs. Others charge for warm pool capacity even when no agent inside that pool is active. Storage billing should mean storage, and only storage, with no hidden compute floor quietly accruing in the background while the agent sleeps.

Sources

  1. AI Agent Sandbox Costs: How Scale-to-Zero Fixes Them
  2. The 2026 Cloud Storage Reset: Why AWS, Azure, and Google Just Made Your Storage Bill Bigger
  3. AI Agent Cost Control: 2026's Shift to Session Caps
  4. Rethinking AI Cloud Infrastructure for Agentic Serving Systems with the Aries Experimentation Framework
  5. Top 5 AI Agent Hosting Platforms 2026 · PandaStack

More in Agent Hosting Economics