Est.

Why Serverless Functions Are the Wrong Primitive for Persistent Agents

Persistent agents need long-lived environments, not stateless functions rebuilt on every call.

Staff Writer · · 11 min read
Cover illustration for “Why Serverless Functions Are the Wrong Primitive for Persistent Agents”
Agent Hosting Economics · October 5, 2026 · 11 min read · 2,570 words

Serverless functions were built for a narrow and well-understood job: run a short piece of code in response to an event, return a result, and disappear. The model treats every invocation as a self-contained unit of work, with nothing expected to carry over once the function finishes. That design was deliberate rather than an oversight. Statelessness is what makes horizontal scaling trivial, billing clean, and operations simple for the workloads serverless was built to handle. A webhook handler, an image resizer, a piece of API glue code, a batch transform job: these tasks are genuinely self-contained. They carry everything they need inside a single request and owe nothing to whatever ran before them.

Every serverless platform rests on three assumptions, stated or not. First, each invocation is independent of every other invocation. Second, no local state survives between calls. Third, the environment itself can be torn down and rebuilt at any time, without warning and without consequence. None of these are flaws. They are load-bearing features, the exact properties that let a cloud provider run millions of functions across shared hardware without ceremony. The trouble starts only when a workload comes along that violates all three assumptions at once, because then the very things that make serverless cheap and scalable start working directly against what the workload needs to do.

What agents do across a session

Agents do not behave like functions. They accumulate state, return to work they started earlier, and operate as one continuous thread stretched across many steps and often many sessions. Consider a coding agent tasked with fixing a bug in a large codebase: it clones the repository, installs dependencies, and boots a development server. The user then asks a follow-up question twenty minutes later. The agent needs to resume inside that same workspace, with the same files on disk and the same server still running, not rebuild the whole environment from scratch because a new question came in.

The same pattern appears in data analysis. An agent loads a multi-gigabyte dataset, cleans it, builds a feature pipeline, and produces a chart. When the user asks a follow-up question about that chart, the sensible path is to continue from the prepared state already sitting in memory, not to re-upload the dataset and re-run every processing step from the beginning. Across a session, agents carry forward tool outputs, intermediate files, running processes, credentials, and a memory of the conversation itself. Every one of these is an artifact of prior work that the next step depends on.

Multi-turn workflows make the dependency chain longer with each exchange. A single reasoning loop might call several tools, write intermediate files, or spawn sub-agents to handle pieces of the task in parallel, and each new step adds another link to a chain that, if broken, forces a full rebuild from the start. Agents also spend real time waiting: for a tool to respond, for a human to approve an action, for a timer to elapse. That rhythm, long stretches of idleness interrupted by short bursts of heavy compute, looks nothing like the clean, synchronous call-and-return pattern a function was built to handle.

The three serverless assumptions agents break, one by one

Diagram: Why Serverless Assumptions Break for Agents. Visualizes: Show three paired contrasts — each serverless assumption alongside the agent reality that violates it — as a vertical stepped list.

Each of serverless's three founding assumptions turns out to be exactly the wrong contract for a workload that accumulates state and keeps returning to prior work.

The first assumption broken is independence between invocations. Agents are not independent calls that happen to arrive one after another. Each turn depends directly on what earlier turns produced, so forcing independence onto an agent means every single invocation pays the full cost of rebuilding context: re-cloning the repository, reinstalling dependencies, reloading the dataset. That cost is not a minor inconvenience. Cloning a large enterprise repository alone can take over two minutes before the agent has done any actual analysis, and multi-gigabyte repositories push that figure higher still. Call this the startup tax: an ephemeral agent pays it on every invocation, and the tax compounds across every turn in a session that might run for hours.

The second assumption broken is that no local state survives between calls. Agents build up state that is expensive to recreate: files on disk, datasets already loaded into memory, development servers already running, sessions already authenticated. A serverless platform destroys all of it the moment the invocation ends. Teams that run into this wall tend to compensate by pushing state out to an external database or object store, which keeps the state alive but does not make the underlying mismatch disappear, only moves it elsewhere. A related failure occurs when there's no persistent memory to fall back on: developers stuff prior context directly into the prompt sent to the model, which inflates token costs on every call and degrades the quality of the response as the context window fills with material the agent has already seen once.

The third assumption broken is that the environment can be torn down at any time. Serverless platforms reserve the right to terminate and replace an environment whenever it suits them, which is a reasonable policy for a stateless function and a disaster for an agent in the middle of a task. A function killed mid-execution loses its place in a multi-step reasoning loop, and there is no way back unless the developer has built checkpoint logic by hand. Real-time and voice agents make this failure the most visible: cold start latency for production language models can run anywhere from several seconds to tens of seconds, and that number has to fit inside a conversational pipeline budget of well under a second if the exchange is going to feel natural. It doesn't come close.

Why workarounds relocate the mismatch instead of fixing it

The standard responses to serverless's statelessness, an external state store, a stuffed context window, a checkpoint framework, each move the cost to a different part of the system without resolving the underlying conflict between invocation-scoped compute and session-scoped work.

External state stores, whether a database, an object store, or something like Redis, keep the state alive by serializing it out and deserializing it back in on every invocation. The state survives, but the round trip has a price: retrieval latency paid on every call, and the real possibility that a complex in-memory structure loses some of its fidelity each time it gets flattened for storage and reconstructed afterward. Context window stuffing is the path most developers reach for first, but with no persistent memory to rely on, prior conversation, tool outputs, and intermediate results all get re-injected into every prompt, burning tokens on information the model has already processed once and making this approach the most expensive at scale.

Checkpoint and durable-execution frameworks, things like Step Functions or similar orchestration middleware, add a different kind of cost: operational complexity that now belongs to the developer. Retry logic, serialization, and resumption plumbing get tangled up with the actual agent logic, turning what should be a property of the infrastructure into a problem every team has to solve for itself. All three approaches share the same shape: they accept a primitive that was wrong for the job and build scaffolding around its limitations, trading added latency, added cost, or added complexity for a problem that a different execution model would not have created.

None of this makes serverless the wrong choice everywhere. One-shot tasks that call an external API and return, stateless validation checks, batch transforms where each item in the batch is independent of every other item: these workloads gain nothing from persistent state, and serverless remains the right tool for them. The mismatch only appears once a workload has session scope, once there is a "before" that the "after" needs to remember.

The right execution model: persistent, isolated compute

An execution model that actually matches how agents behave needs three properties at once: it has to persist, it has to isolate, and it has to rest cheaply when nobody is using it. Persistence means the environment survives between invocations, so the repository is already cloned, the dependencies are already installed, and the dataset is already sitting in memory. Instead of a multi-minute rebuild, the agent resumes in milliseconds.

Isolation has to happen at the level of the virtual machine, not the container, because this is a correctness requirement for multi-tenant systems, not an optional security upgrade. Containers share a host kernel. A micro-VM gives each agent its own guest kernel inside a hardware-virtualized boundary, so an attacker who compromises one agent still has to escape both the guest kernel and the hypervisor before reaching the host or another tenant's data. Firecracker, the open-source Virtual Machine Monitor written in Rust, is built around exactly this guarantee: each micro-VM runs its own kernel, and reaching the host means escaping both that guest kernel and the KVM hypervisor underneath it. This is why reusing a single persistent VM across two different users to save on provisioning cost is not a shortcut worth taking. It collapses the isolation boundary and opens the door to cross-tenant data exposure. The only safe model assigns one VM per trust boundary.

Hibernation is what makes persistence affordable rather than wasteful. Persistent does not have to mean always running. A snapshot-and-restore operation captures the full machine, filesystem and memory together, to storage, then stops the VM. State survives, but it now costs storage rather than compute. On the next request, the agent restores from that snapshot, and the user experiences a warm resume without the operator having paid for a machine that sat idle around the clock. None of this works if the wake-up is slow: a persistent agent that takes several seconds to restore has just reintroduced the cold start it was supposed to eliminate. The resume path has to be fast enough that continuity feels real to the user, and Firecracker's snapshot-restore architecture was built specifically to hit that bar.

The economics of persistent agents at scale: idle cost is not the objection it appears to be

The usual objection to persistent compute is that it must cost more than serverless, since a serverless function only gets billed while it runs and a VM sounds like it runs all the time. That objection collapses once hibernation enters the picture, because most agents spend most of their time idle, and in a snapshot-restore model, an idle agent is paying for storage, not compute.

An always-on persistent VM is genuinely expensive to run around the clock. A hibernated persistent VM is not, because the two are paying for entirely different things: one pays for uptime, the other pays for the state captured in a snapshot sitting in storage. As a fleet of agents grows, this distinction only sharpens. A serverless billing model charges per invocation regardless of how much reconstruction that invocation has to pay for. A hibernated persistent model charges compute only for the sessions that are actually active right now, and charges storage rates for everything else that is dormant. The operator ends up paying for concurrency, not for the total number of customers on the books. In a multi-tenant fleet where only a fraction of users are active at any given moment, that distinction is the whole economic argument: the hibernated fleet costs a fraction of what an always-on equivalent would, and that fraction keeps shrinking as the ratio of dormant to active users grows.

The same snapshot mechanism opens up a second capability almost as a side effect. Branching from a single warm snapshot to explore several possible next steps in parallel, the kind of tree-of-thought reasoning agent systems increasingly rely on, costs next to nothing per branch. Each branch is still a fully isolated machine, and a branch that turns out to be a dead end can be discarded at the cost of simply killing a process.

The per-user agent fleet as the reference architecture for agent-native products

One micro-VM per user emerges as the pattern: that user's identity is injected at launch, their conversation memory is held in RAM, and their session is hibernated down to storage cost the moment they go quiet. That pattern turns a user-scoped, persistent agent into something a product can actually be built on, rather than an infrastructure expense a company has to justify separately.

The per-user model lines up with how people actually experience agent products. A user's agent remembers the files it worked on, the preferences it picked up, the questions already answered. For the per-user model, continuity across sessions is the product itself, not just an engineering detail. Managing a fleet built this way turns into a scaling question rather than a provisioning headache: launching a new user's agent is a single operation, and the fleet's total cost tracks active concurrency rather than the raw count of registered users sitting in a database somewhere.

Security falls out of the architecture directly rather than depending on a configuration setting someone has to remember to apply. Each user's VM boundary is the isolation unit, so a breach in one tenant's session requires escaping the hypervisor itself, not just slipping past a software namespace that a container shares with its neighbors.

Cloud providers have started building toward exactly this model. In June 2026, AWS launched Lambda MicroVMs, described as a new serverless compute primitive within AWS Lambda that lets developers run code generated by users or AI in isolated, stateful execution environments. The launch includes extended-duration sessions, startup from a Firecracker snapshot, and per-session allocation of compute and disk. A major cloud provider is now confirming, through a shipped product, that stateful, VM-isolated, snapshot-started execution is the right model for agent workloads, which is the same direction this argument has been building toward from the start.

What to look for in infrastructure for persistent agents

Evaluating infrastructure for agent workloads comes down to a short list of concrete questions, each one tracing back to a mismatch described earlier in this piece.

Persistence is the first question to ask: does the environment keep the filesystem and memory intact between invocations, or does the developer end up owning state serialization by hand? The answer decides whether the infrastructure fits how an agent actually behaves or fights it on every turn.

Isolation level determines how far a breach can spread, so it is not a detail to gloss over. Container-level isolation, built on a kernel shared with every other tenant on the box, is not the same guarantee as VM-level isolation, where each agent runs its own dedicated kernel. For a multi-tenant agent fleet handling real user data, that difference is the one that determines whether a single compromised session stays contained or spreads.

Wake latency belongs on the same checklist: a resume that takes several seconds quietly brings back the delay of starting cold that persistence was supposed to solve. The idle cost model deserves equal scrutiny, since a platform that charges compute rates for dormant sessions has not actually solved the economics, no matter what it calls itself. And deployment simplicity, how much plumbing a team has to build on top of the platform just to get checkpointing, retries, and resumption working, is often the clearest signal of whether a piece of infrastructure was designed around the agent's behavior from the start or adapted after the fact. The direction the industry is moving, visible now in a major cloud provider's own product launch, points toward persistent, isolated, snapshot-based compute as the standard agents were missing from the beginning.

Sources

  1. Firecracker

More in Agent Hosting Economics