Per-Agent Pricing vs Shared Infrastructure Cost Models
Infrastructure choices determine which pricing model your business can actually afford to sustain.

Most vendors still build a pricing page the way they always have: run a demo, see what the buyer will tolerate, write the number down. The order that actually holds up is the reverse of that. The cost structure a team can sustainably offer is set by how its agents are deployed, isolated, and kept idle long before anyone opens a spreadsheet to model margins. Picking between per-seat, per-resolution, or usage-based billing looks like a go-to-market decision, but that choice depends on whether the infrastructure underneath can support the billing granularity the model requires. A team running agents on shared infrastructure cannot offer true per-agent billing without absorbing unpredictable cross-tenant costs, and a team running isolated per-agent VMs cannot spread idle overhead across a pool to flatten its unit economics. Each path forecloses the other, and that's the architectural fact the rest of this piece works through.
How the traditional per-seat model broke under agent workloads
Per-seat pricing carries a built-in assumption: value scales with the number of humans using the software. That assumption collapses the moment the software finishes the work itself rather than helping a person do it faster, because an agent that resolves thousands of conversations a month doesn't add a single seat to the invoice, since seats count licensed humans, not completed work.
The mismatch appears directly in the billing line. A company running an AI agent that handles the bulk of its inbound support volume still pays for every human seat on the contract, whether the AI resolved a handful of conversations that month or tens of thousands. The fee doesn't move with the work, so the vendor gets paid the same amount whether the agent is nearly idle or running at full capacity, and the buyer has no lever to pull except adding or removing humans, which has nothing to do with what the agent is actually doing. A per-seat contract rewards the customer for keeping more humans on staff, putting the vendor's incentives at odds with the customer's, since the whole point of deploying an agent is to reduce reliance on human staff.
The market has already registered this break. Seat-based pricing fell sharply among SaaS companies in a single year, while hybrid pricing jumped to become the dominant structure, per DEV Community's 2026 pricing analysis; that shift reflects a structural incompatibility between seat-counting and agent workloads rather than a change in buyer taste. The underlying cost structure makes the same point from the supply side: agent compute cost is not flat. A single query can cost an order of magnitude more or less than another depending on input complexity, so a fixed per-seat fee simultaneously overcharges the light user and undercharges the heavy one. Per-seat pricing is wrong because it bills on a unit, the human seat, that has stopped correlating with the thing actually generating cost and value, which is agent activity.
The three billing models that have emerged to replace per-seat
Three models have emerged to take per-seat's place: per-conversation (also called per-ticket), per-resolution, and outcome-based billing. Treating these as a neutral menu, where a team just picks whichever sounds most palatable to its buyers, ignores what each one actually requires underneath. Every one of the three encodes a different bet about how the agents generating the bill are deployed, and a team that adopts a billing model without checking whether its infrastructure can back that bet is setting itself up for a margin surprise.
Per-conversation billing charges a fee for every inbound interaction the agent touches, resolved or not, typically somewhere between low cents and around a dollar per inbound. Vendor revenue under this model rises with inbound volume regardless of whether the agent actually solves anything, which rewards a noisy inbox and penalizes a team that successfully reduces inbound through better product design. The infrastructure bet here is modest: because the billing unit is the conversation and not the agent instance, shared compute pools work fine, and paying for per-agent isolation would add cost without adding any billing precision that per-conversation pricing could use.
Per-resolution billing only charges when the agent closes a job without handing it to a human, and nothing accrues on conversations that escalate or fail. Quickchat AI's April 2026 pricing comparison lists a published per-resolution rate of $0.50 for Quickchat AI Enterprise and for HubSpot Customer Agent, with competing platforms priced meaningfully higher on the same unit. That's a real alignment benefit: vendor revenue tracks the outcome the buyer actually cares about. But "resolution" is defined by the vendor, not by any shared standard, so two vendors quoting an identical rate can end up billing very differently for what looks like the same work. The infrastructure bet is heavier than per-conversation's: the platform has to attribute a completed outcome to a specific agent run, which is straightforward on shared infrastructure in principle but only becomes auditable, and therefore trustworthy to a buyer who wants to check the vendor's math, when agents are isolated enough to carry a traceable identity.
Outcome-based billing goes a step further, charging for an attributable business result, a qualified lead, a booked appointment, a payment collected, and nothing when there's no result. Outcome-based billing only works when both sides agree in advance on what counts as the outcome and can measure it without dispute, because a vague definition will drift in whichever direction favors the party writing the contract, which makes measurement the chief obstacle to this model. It requires per-agent identity, persistent state to track a result across a multi-step workflow, and audit logging detailed enough to resolve a disagreement after the fact.
None of the three is categorically the right or wrong choice. The mistake is procedural: a team can adopt one of them without first confirming the deployment underneath can actually deliver the attribution, state tracking, or audit trail the model assumes. That's likely why hybrid pricing, a base subscription combined with usage and outcome components, has become the most common structure in the market and keeps growing. No single model covers the full range of workload shapes a vendor is likely to encounter across its customer base.
Why idle time is the dominant cost variable in any agent fleet
Each of the three models above is a hidden bet on infrastructure, assuming the underlying deployment matches its particular assumptions. Agent workloads are event-driven. An agent spends most of its existence waiting rather than working. Idle cost, not active compute cost, is the dominant cost variable across most fleets, and the billing model a team can sustain turns out to be largely a function of how well its infrastructure handles that idle overhead.
Stated concretely: a per-agent billing model only works economically if the team can stop paying for compute the moment an agent goes dormant. Without that, the cost of running a fleet of thousands of agents accrues continuously, whether those agents are doing anything or not, and the vendor either eats the loss or passes it through as a padded rate. Shared-infrastructure models sidestep the problem by pooling compute across many tenants so idle time from one customer gets absorbed by active time from another, but that pooling comes at the cost of isolation and billing granularity, and that tradeoff is what determines which pricing structures are even available downstream.
The architectural answer to the idle problem is the snapshot-restore pattern: an agent's filesystem and memory state get saved to a snapshot the moment it goes idle, billing drops to storage cost alone, and the agent restores to a live state in well under a second when the next request comes in.
If memory is handled carelessly, it makes the agent's idle periods even more expensive to recover from. Epoch AI's analysis finds that inference costs have fallen dramatically per year, though the rate of decline varies a great deal by workload. That trend doesn't eliminate the cost of rebuilding an agent's context from scratch on every wake. Most of the input tokens an agent burns through in production are system prompt overhead rather than conversation history. A persistent, per-agent memory store isn't a nice-to-have optimization at scale but a financial necessity, since reconstructing that context on every single wake is itself an expensive operation.
The pricing implication follows directly. That's the reason a fair number of pricing pages labeled "per-agent" turn out, on close inspection of their rate structure, to be flat subscriptions wearing different branding.
How microVM isolation changes the economics of per-agent billing
MicroVM isolation is the prerequisite that makes true per-agent billing granularity technically possible and makes per-tenant cost attribution accurate in the first place.
Firecracker, the microVM technology behind this pattern, runs one VMM process per microVM rather than a single daemon managing every VM on a host. That architectural choice is the foundation of true per-tenant billing granularity, because cost attribution can follow the process boundary directly instead of being estimated or apportioned after the fact. Firecracker boots in roughly 125 milliseconds with about 5MB of memory overhead; a single bare-metal host can run thousands of instances at once. Per-agent isolation, in other words, does not require per-agent hardware, and that's what makes the economics work at scale rather than only in a lab demo. Each microVM also runs its own guest kernel, so one agent can install packages, run arbitrary processes, and even break its own environment without touching any other agent on the same host. The blast radius of a misbehaving agent stops at the VM boundary.
That last point is where the security argument and the billing argument converge, and where shared infrastructure runs into a limit it can't fully answer. AI agents generate unpredictable code by nature, and a shared kernel means one misbehaving agent or tenant can affect the entire pool running on it. The risk extends beyond a surprise on the cost side to a security blast radius, and organizations in regulated industries are increasingly treating isolation as a hard requirement for any agent deployment rather than a preference to weigh against cost. Forrester's AI Infrastructure Survey from 2026 found that roughly half of Fortune 500 companies running AI agent workloads now use Firecracker-backed sandboxes, a sign that the market is resolving isolation toward microVM isolation even in deployments that started out on shared infrastructure.
Combining snapshot-restore with microVM isolation collapses the tradeoff that seemed fundamental a few paragraphs ago, producing isolation without idle compute cost: the VM can be snapshotted to storage and restored in under a second, with billing covering only storage between sessions. One sandbox provider's pricing, listed in Infrabase.ai's agents pricing comparison, offers a free tier with introductory credits followed by a paid tier billed per second on compute, with billing stopping entirely when the sandbox is paused. A separate persistent, hardware-isolated Linux environment built on Firecracker microVMs, launched for AI agents in January 2026, creates new instances in seconds, checkpoints and restores in about one second, idles automatically when inactive, and carries a large NVMe filesystem that survives indefinitely between sessions, with compute billing stopping while the environment is idle. Neither case is cited here as a recommendation over the other. Both demonstrate the same mechanism from different angles: isolation and idle-cost elimination aren't in tension once the infrastructure is built to support both at once.
What the per-user agent fleet pattern demands from the infrastructure layer
Deploying one agent per user, rather than one shared agent serving a whole customer base, is shaping up as the right primitive for the next generation of agent-native software. It also exposes every infrastructure weakness that a shared pool or a stateless deployment can paper over at small scale, because the requirements only get harder as the fleet grows, not easier.
Isolation requirements multiply first. In a per-user fleet, each agent needs to be isolated not just from outside threats but from every other user's agent running on the same host. Credential leakage, filesystem cross-contamination, and a shared-kernel exploit all turn into user-data incidents the moment they occur at multi-tenant scale. That's why per-user agent deployment effectively requires microVM-level isolation: a container boundary isn't enough once agents are executing arbitrary code while holding a given user's credentials.
Persistent state is the second requirement, and it's just as non-negotiable. A per-user agent has to carry memory, credentials, and workflow state across sessions. An agent that loses its context every time it wakes is a stateless function wearing per-user branding. State here means a snapshot of everything the agent currently knows about the task in front of it: what step it's on, what the last tool call returned, what variables it's tracking. None of that survives by accident. It has to be persisted on purpose.
Identity and authorization make up the third requirement, and for many product teams building on shared infrastructure, this requirement is what stalls launches. Each user has to authorize their own account inside the product, under the builder's brand, with tokens kept server-side and refreshed automatically rather than handled manually per session. AWS Bedrock AgentCore's on-behalf-of token exchange, where an agent swaps an inbound user access token for a new, scoped token aimed at a specific downstream resource, has become something close to the canonical pattern for running per-user agent fleets on shared infrastructure, and it shows what identity-binding an infrastructure layer has to support. Bloomberg's engineering team has described a similar gap in its own deployments, what it calls the "productionization gap," the lag between a GenAI demo that works and an application that's actually deployable and compliant. Bloomberg closed that gap using identity-aware, multi-tenant MCP servers, treating prompts and toolchains as configuration rather than code, which cut its experimentation cycle from weeks down to minutes.
All three requirements point to the same economic conclusion. Per-user agent deployment is only viable if idle agents pay storage cost rather than compute cost. Without that, fleet cost scales with user count at every hour of the day rather than with actual usage, and the unit economics break down at any scale past a small pilot. The pricing model a team can offer its own end users, per-user, flat subscription, outcome-based, whatever shape it takes, is bounded by whether the infrastructure underneath can support idle-cost optimization for every individual agent instance in the fleet.
The strongest objection to per-agent pricing, and shared infrastructure's incomplete answer to it
The strongest case against per-agent pricing is a cost-volatility argument, and it deserves to be taken seriously rather than waved off. A team billing per agent is exposed to every spike in usage pattern across its customer base: a single enterprise customer that suddenly runs ten thousand agents for a batch job creates a compute bill that has to be covered somewhere before the vendor ever collects a cent from the customer who caused it. Shared infrastructure is the conventional answer to that volatility, because pooling compute across many tenants smooths out exactly this kind of spike, letting idle capacity from quiet customers absorb the burst from active ones.
Pooling compute across tenants does smooth out cost volatility, but it only resolves that problem by reintroducing the isolation and attribution problems the earlier sections laid out. A vendor can have precise, auditable, per-agent billing, or a vendor can have idle-cost smoothing through pooled shared infrastructure. Snapshot-restore combined with microVM isolation is the only approach on the table that gets close to both at once, isolation without a continuously accruing idle bill, rather than a true third option. Any team evaluating its own pricing page needs to ask whether the infrastructure underneath actually supports the billing model on offer, or whether the billing model is quietly borrowing against a tradeoff nobody has resolved yet.
Sources
- AI Agent Pricing Models 2026: Per-Resolution vs Per-Seat Compared
- AI Agent Pricing in 2026: Per-Seat, Per-Resolution, or Outcome? - DEV Community
- AI Agent Pay-Per-Use Pricing
- Agents Pricing: 31 Providers Compared (2026) at Infrabase.ai
- Firecracker for AI Agent Sandboxes: Benefits, Limits, and Evaluation Questions - Novita
- AI Agent Code Execution Sandboxes: Isolation from Containers to MicroVMs
