Est.

Metered vs Flat-Rate Pricing Tradeoffs for Agent Infrastructure

Metered pricing fits agent workloads; flat rates fit deployment slots.

Senior Contributor · · 10 min read
Cover illustration for “Metered vs Flat-Rate Pricing Tradeoffs for Agent Infrastructure”
Agent Hosting Economics · October 7, 2026 · 10 min read · 2,193 words

Metered versus flat-rate pricing for agent infrastructure is a structural question with a right answer for every given workload, and the answer depends on how that workload actually consumes resources over time, not on which model looks simpler on a pricing page.

The flat-rate model that worked for SaaS breaks under agent workloads

Flat-rate and per-seat pricing worked for SaaS for a specific reason: human usage stayed predictable enough that the cost to serve any one customer looked roughly like the cost to serve any other. A seat costs about the same whether the person behind it logs in once a day or fifty times. That stability let a vendor set one price, apply it across the customer base, and expect margins to hold steady no matter which accounts were active in a given month.

Agent workloads don't behave that way. An agent's cost to serve can swing by orders of magnitude depending on the task it happens to run, and that swing can happen inside a single account from one day to the next. Apify's analysis of agent pricing models states the problem directly: an agent can sit doing nothing for three days, then run a hundred jobs in an hour. The cost of serving it isn't steady across that window, so pricing it as if it were means picking one of two losing outcomes: overcharging the customer during the idle days, or undercharging during the burst. The same unevenness appears inside a single agent's own task list: a support agent resolving one difficult ticket can burn far more tokens than the same agent answering a one-line question minutes later, so the variance is present within one agent's normal working day and not just across a fleet of agents.

Agents also introduce a failure mode that flat-rate SaaS billing never had to account for: the retry loop. Capping usage sounds like an obvious fix, but a cap without visibility into which agents are burning what tells a team nothing about where the cost is coming from, and a hard cap enforced mid-task simply breaks the agent. The retry loop is itself one of metered billing's blind spots, which this piece returns to later. For now, the point is narrower: the assumption that made flat pricing work for SaaS, that cost to serve stays roughly constant, does not hold for agents, and every pricing decision that follows has to start from that fact.

The two dimensions every agent infrastructure bill is made of

Diagram: The Two Layers of Every Agent Infrastructure Bill. Visualizes: Visualize the structural split between two cost types in agent infrastructure billing.

Before you compare pricing models, it helps to separate what's actually being priced. Every agent infrastructure bill is made of two distinct cost types, and you need to sort each cost onto the right side of that line. Apify's analysis states the rule directly: costs that are predictable belong on a subscription, and costs that vary with what the agent actually does belong on a meter.

The predictable layer covers the capacity a provider has to keep ready regardless of whether the agent uses it: compute reserved for the agent's deployment, storage allocated to it, the number of concurrent runs the plan supports. A provider knows roughly what it costs to hold that capacity available, so it can map a flat subscription fee cleanly onto it. The variable layer covers everything that depends on what the agent decides to do once it's running: LLM calls, tool invocations, external API requests, active execution time. None of that volume is determined by which plan the customer picked. It's determined by the task in front of the agent. Bundling it into a flat fee guarantees the provider will guess wrong in one direction or the other for most customers.

Lago's explanation of metered billing draws the same boundary from the billing side rather than the infrastructure side: subscription billing gives customers predictable costs, so it remains the default for products with steady, bounded usage, while metered billing fits products where cost genuinely scales with consumption, naming infrastructure, AI inference, and API platforms as the clear cases. Most real deployments need both layers running at once: a flat base covering the infrastructure floor, and a meter covering the consumption that varies underneath it. The choice was never flat versus metered as an abstract preference; it's which specific cost types get assigned to which layer.

What metered billing measures

Metered billing is precise about the thing it measures, and that precision is also its limit: whatever falls outside the metering boundary doesn't appear on the bill at all, however large it turns out to be. Lago's breakdown of how metered billing works describes a four-step pipeline that underlies every usage-based invoice: usage gets captured as events, duplicate events get filtered out so a retried event doesn't bill twice, the filtered events get aggregated into a billable metric, and a rate gets applied to produce the invoice.

The aggregation rule a vendor chooses at that third step changes what the customer is actually being charged for, even when the underlying usage is identical. COUNT bills for raw activity. UNIQUE COUNT bills for distinct identities. SUM bills for volume. MAX bills for peak commitment. The same stream of agent activity produces a very different invoice depending on which of these four rules sits behind the meter, so the choice of aggregation method is itself a pricing decision, not a neutral technical detail.

Claude Managed Agents shows what a metered model looks like when the provider draws a deliberate line around what counts as usage. A scope decision that helps one workload shape can hurt another, so you need to know which shape describes your deployment, not just the headline rate.

When flat-rate infrastructure pricing is structurally correct

Flat-rate pricing for agent infrastructure isn't just a fallback for teams that want to dodge metering complexity. If the thing being priced is the deployment slot itself rather than the work an agent does inside that slot, this is the structurally correct model. Persistent, always-on agents, scheduled scrapers, queue workers, an agent that holds a continuous identity in a chat channel, all occupy their deployment slot continuously regardless of how much task volume passes through them in a given week. The scarce resource in that case is the slot, not the throughput, and a flat per-agent fee is the pricing shape that matches a resource whose cost barely moves with usage.

The agent pricing comparison analysis frames the decision as a rule of thumb: pick a pricing unit that tracks what can actually be forecast. Maritime's own plan structure reflects this logic directly, pricing a flat monthly rate per agent for the VM, the persistent disk, and the network allocation that keep an agent's identity available, while treating execution and tool invocations as the natural candidates for metering layered on top. Separating the two cleanly in the architecture, rather than trying to sort them out after the fact in a billing system, is what keeps the model from becoming ambiguous later.

The distinction that matters most here is between flat-rate on the infrastructure slot and flat-rate on the inference happening inside it. That's why a flat rate on the deployment slot holds up: the provider's cost to keep the slot available doesn't move much with what happens inside it. Token and compute consumption inside that slot can vary by orders of magnitude between two customers paying the identical fee, so a flat rate on the model inference running inside it doesn't hold up at scale.

Workload shape, duty cycle, burstiness, and idle ratio determine which model fits

The duty cycle, the fraction of clock time the agent spends actually working rather than waiting for the next task, predicts which pricing model will cost less and which will produce runaway exposure for a given agent.

A low-duty-cycle agent is bursty and mostly idle, so it fits consumption billing. Infrastructure market data points to a crossover somewhere in the sixty-to-eighty percent utilization range, below which paying for active time tends to beat paying for reserved capacity, and above which the reverse holds.

Burstiness is a separate variable from duty cycle and has to be checked on its own terms. Flat-rate capacity reserved to cover that peak sits expensive and unused for most of the month if the agent only hits peak load occasionally.

GitHub Copilot's migration in June 2026 is a documented instance of workload shape forcing a change in pricing model. Lago points to this as a case where agentic coding sessions consumed far more compute than a flat per-request allotment could reasonably account for, which made the flat structure unsustainable once agentic use grew heavy enough. The shift wasn't a pricing experiment; it was a correction once the workload outgrew the model built for a lighter one.

The practical step before choosing a pricing model is to measure, or at least estimate, three numbers: average duty cycle, peak concurrency, and idle ratio. Without those three figures in hand, comparing the sticker price of a flat plan against a metered one tells a team almost nothing useful, because the same headline rate can be a bargain for one workload shape and a loss for another.

Diagram: Duty Cycle Determines Which Model Wins. Visualizes: Show how agent utilization (duty cycle) predicts which pricing model costs less.

Snapshot-restore scheduling makes idle time cheap under either model

Long-lived agents spend most of their existence idle, as they wait on the next message, the next scheduled run, the next human approval. Conventional deployment keeps a full virtual machine resident for each one whether or not it's doing anything, and that unused resident compute is the real cost that flat-rate pricing for idle agents has to absorb somewhere.

Snapshot-restore scheduling changes that math by pausing an idle agent to disk and restoring it the moment the next inbound task arrives, converting idle cost from compute cost into storage cost. The engineering difficulty sits in snapshot size: a sandbox with a large RAM and disk footprint creates a large volume of data to move in and out of storage on every pause and resume, which becomes a latency problem unless the system is built to snapshot only the chunks that changed. Content-addressable, copy-on-write storage keeps that transfer small by moving just the dirty data.

Google's Agent Substrate shows what this enables at scale: it multiplexes a large number of stateful agent sessions onto a comparatively small number of physical Kubernetes pods, holding full state intact across hibernation cycles while running at a high oversubscription ratio. Google Cloud Blog reports that suspend-and-resume scheduling with GKE Agent Sandbox cut per-agent cost by up to 75 percent. Without a mechanism like this, a flat rate applied across thousands of idle agents means paying full compute for every one of them. With it, idle agents pay largely for storage, a cost that can run an order of magnitude below the compute they'd otherwise occupy.

Maritime's infrastructure is built around the same principle, using Firecracker micro-VMs with snapshot-and-restore scheduling so agents sleep between tasks and wake in about a second. Running each agent in its own VM with continuous availability lets a platform show a customer exactly which agents are consuming resources and when, drawing a clearer line between infrastructure cost and execution cost than a shared, stateless environment typically allows. Metering schemes that exempt idle time, like the session-runtime model Claude Managed Agents uses, solve a related problem from the billing side. They help the bursty case well, but they leave a continuously running agent accumulating runtime charges before a single token gets counted. Snapshot-restore scheduling addresses the same gap at the infrastructure layer instead, so the billing model lines up with actual resource consumption regardless of which pricing structure sits on top.

The hidden costs that neither metered nor flat-rate pricing makes visible

A meaningful share of what an agent deployment actually costs to run falls outside what either flat-rate or metered pricing puts on the invoice. The agent pricing comparison analysis names the specific items that tend to move a real bill without ever appearing on a plan card: model retries, context growth across long-running sessions, devbox activity, memory writes, and outbound API calls. None of these appear in the headline number a customer compares when choosing a plan, yet all of them draw on real infrastructure.

Plan language carries its own ambiguity here too. "Unlimited agents" means little if concurrency or runtime capacity is capped somewhere else in the fine print, so the plan description and the effective limit a customer actually operates under can diverge substantially. Retry loops are the sharpest version of this problem under a metered model specifically: an agent retrying a failed tool call burns active execution time and tokens while producing no useful output, and the billing system has no way to distinguish that from productive work. The charge is real even though the result isn't.

Multi-tenant fleets add a separate complication once snapshot-restore enters the picture. When many agents are multiplexed onto shared infrastructure that pauses and resumes them on demand, the standard assumption behind cost attribution, that one process maps to the team or customer it serves, stops holding. Pricing a given agent correctly requires knowing its duty cycle and the layer each cost belongs to, and it also requires tooling built for an architecture where agents genuinely come and go.

More in Agent Hosting Economics