Designing a Hybrid Subscription and Token Pricing Model That Doesn't Lose Money
Blending a flat subscription fee with usage-based token pricing is a reasonable way to balance predictable revenue against variable cost, but the blend creates its own failure modes that a pure subscription or pure usage model doesn't have. The most common one: a small share of heavy users consuming far more than their subscription covers, quietly eating into margin on every account like them.
Here's where that architecture tends to break, and how to check whether it's already happening in your own pricing.
The core vulnerability: usage that outpaces the subscription
A flat subscription fee is calculated against an assumed average usage level. Any customer using meaningfully more than that average is being subsidized by the ones using less, and if your product attracts power users disproportionately, or a feature encourages heavier usage than the pricing model assumed, that subsidy can be large enough to turn a profitable-looking plan into a loss leader for its heaviest cohort.
This is easy to miss because total revenue can look healthy even while a specific tier is quietly unprofitable, since strong performance elsewhere in the business masks it. The gap only becomes visible when you look at unit economics tier by tier and account by account, rather than at the company-wide top line.
Auditing your own model for this gap
Pull actual token usage by account, calculate the token cost each account is generating, and compare it against what that account is paying. Plot accounts by usage and look at where the cost curve crosses the revenue line. If a meaningful share of accounts sit above that crossing point, especially if that share is growing, your pricing has this vulnerability today, not hypothetically.
Segment the plot by plan tier as well as by account, since the gap often concentrates in one specific tier, commonly the entry tier priced to win new customers, rather than being spread evenly across your whole customer base.
Audit and fix the gap in this order:
- Pull actual token usage by account from your billing or usage data, rather than working from the average you assumed when pricing.
- Calculate the token cost each account generates and compare it against what that account pays.
- Plot accounts by usage and find where the cost curve crosses the revenue line.
- Set the included allowance against a percentile of actual usage rather than the average.
- Price overage tokens with real margin over your own cost, not at or near cost.
- Move existing customers onto changes through a grandfathered or clearly explained transition path.
Where the included usage allowance should actually be set
Set the included allowance in a subscription tier against a percentile of actual usage, not the average, since a plan built around the average will always be underwater for the heavier half of its users. A common approach is pricing the included allowance to cover a large majority of accounts at that tier comfortably, with overage pricing designed to make the remaining heavy accounts still profitable rather than subsidized.
Overage pricing that actually protects margin
Overage tokens should be priced with real margin over your own cost, not at cost or near it, since overage is exactly the usage that was outside what the subscription was priced to cover. A thin or non-existent overage margin defeats the purpose of having a usage-based component at all: it's there specifically to keep heavy usage from eroding the economics of the plan.
Revisit the overage rate whenever your underlying model cost changes, since a rate set to protect margin against last year's per-token cost can quietly stop doing that job if your cost structure shifts and nobody rechecks the rate against it. Treat it as a rate that needs periodic maintenance, the same way you'd revisit a foreign exchange assumption baked into a contract.
Communicating changes without a customer backlash
If the audit finds a real gap, fixing it usually means adjusting included allowances or overage rates for new customers, and handling existing customers on a grandfathered or clearly communicated transition path. Customers respond far better to a clearly explained shift, especially one framed around usage tiers rather than a blanket price increase, than to a change that looks like the number changed with no explanation attached.
A worked look at where one tier quietly loses money
Say your entry tier includes a fixed monthly token allowance priced to comfortably cover most accounts, but a specific use case, batch document processing rather than the occasional chat query the tier was designed around, pulls a small slice of accounts far past that allowance every month. Those accounts pay overage, but if the overage rate was set closer to a rounding convenience than to your real per-token cost, the tier can still lose money on exactly the accounts driving its heaviest usage.
The fix isn't necessarily to push those accounts to a higher tier, since the use case itself might be perfectly reasonable for an entry-level customer to run occasionally. It's to make sure the overage rate on that specific tier reflects real cost with real margin, checked against the actual usage pattern of the accounts triggering it, rather than a number picked once at launch and left alone.
This is also where segmenting by account type earns its keep: a handful of accounts running an unusual, heavy workflow inside an otherwise healthy tier is a very different problem from the whole tier being mispriced, and the audit needs to distinguish the two before deciding what to change.
What Good Looks Like
Good looks like knowing, by account, whether current usage is profitable under your pricing, not just knowing total revenue against total cost in aggregate.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
How often should we re-audit the pricing model for this gap?
Quarterly is a reasonable cadence for a growing product, since usage patterns can shift meaningfully as new features ship or as your customer mix changes. A model that was well calibrated at launch can develop this gap within a couple of quarters if a popular new feature drives usage up faster than pricing was adjusted for.
Should we cap usage instead of charging overage?
A hard cap avoids the margin risk entirely but tends to frustrate customers more than a clearly priced overage, since it can block a legitimate workflow mid-task. Most products land on overage pricing with a cap as a last resort safeguard, rather than a cap as the primary control.
Is it better to move to pure usage-based pricing instead of a hybrid?
That depends on how much your customers value predictable billing versus paying exactly for what they use. A hybrid model is often the right compromise precisely because customers like predictability, but it only works financially if the included allowance and overage rate are set correctly, which is the whole point of auditing them.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Setting Hard Spend Caps So an AI Agent Can't Run Away with Your Bill
How to design token budgets and circuit breakers for AI agents, so a looping agent or a bad prompt can't turn into a five figure surprise invoice.
How to Recognize Revenue on Usage-Based AI Billing
Token-metered contracts don't fit a simple monthly revenue model. Here's how to apply ASC 606 to prepaid credits, overage fees, minimums, and breakage.
Tracking Token Cost per Active User on a Simple FinOps Dashboard
How to pick the right denominator, the numbers to track weekly, and how to build a token cost per active user dashboard finance actually reads.
Chargebee or Stripe Billing for Negotiated SaaS Deals
A decision guide for B2B SaaS finance teams choosing between Chargebee and Stripe Billing once sales starts negotiating ramp deals and custom terms.
What Evaluating Your AI Agent Actually Costs to Run
See where AI agent evaluation cost comes from: judge-model calls, human review and test set upkeep, with a worked run example and ways to keep spend in check.
Build Your Own AI Inference Cost Model in Three Tabs
How to structure a spreadsheet that turns token usage into a real cost per customer, so you can see GPU and API spend before the invoice arrives.