Reserved GPU Capacity vs Spot: Getting the Accounting Right
Reserved capacity and spot instances solve different problems, and treating them as one undifferentiated "GPU cost" line is how cost allocation goes wrong. A reserved commitment is a fixed cost you've locked in regardless of use; spot capacity is a variable cost that can vanish mid job.
Knowing which is which matters for two separate reasons: picking the right mix, and then allocating the resulting bill to the teams and products actually driving it.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
How the two actually price differently
A reserved instance or committed-use discount trades a term commitment, often one or three years, for a lower effective rate on guaranteed capacity. Spot capacity is unused inventory the cloud provider sells at a steep discount with no guarantee: it can be reclaimed with short notice. The reserved rate is a known number you can put in a budget. The spot rate floats, and the real cost of spot includes the checkpointing and retry work needed to survive interruption.
When reserved commitments pay off
Reserved capacity earns its keep when your baseline usage is predictable: a production inference endpoint serving steady traffic, or a training schedule you run every week regardless of demand. If you can look back six months and see the same GPU hours showing up reliably, that's the volume to commit against. Committing against usage you hope will happen, rather than usage you've already seen, is how idle reserved capacity ends up sitting on the balance sheet as sunk cost.
When spot is the better bet
Spot fits work that can tolerate interruption: batch training jobs with checkpointing, offline evaluation runs, and anything you can resume rather than restart. It does not fit a customer facing inference endpoint where a reclaimed instance means a failed request. The teams that get the most out of spot invest early in checkpoint logic, since the discount is only real savings if a reclaimed job resumes cleanly rather than starting over.
Blending both without losing visibility
Most mature setups run a reserved floor sized to steady state demand, with spot absorbing everything above it. The allocation trap is letting that blend get invisible in the billing export, where reserved and spot usage land in the same line item. Tag workloads by commitment type at the point of launch, not after the invoice arrives, so finance can split the bill by type without archaeology.
Use this checklist to size and allocate a blended GPU setup:
- Set the reserved floor from the GPU hours that show up reliably in your usage history, and commit only against that steady baseline.
- Send interruption tolerant work such as checkpointed batch training and offline evaluation to spot capacity, and keep customer facing inference off it.
- Tag every workload by commitment type when it launches, so reserved and spot usage do not merge into one billing line item.
- Record upfront reserved payments as a prepaid asset amortized over the term, and expense spot and on-demand usage as it is incurred.
- Put the renewal date on the calendar and compare actual usage to the commitment one full billing cycle before it renews.
Booking the entries correctly
Reserved commitments paid upfront are typically a prepaid asset amortized over the term, not a lump expense in the month you pay. Spot and on-demand usage is expensed as incurred. Mixing these up either understates a quarter's true infrastructure cost or overstates it, and both versions make gross margin look wrong to whoever is reading it, including the board.
If you cancel a workload mid-term and stop using a reserved commitment you're still paying for, the remaining unamortized balance doesn't just disappear from the books; it typically still needs to be recognized over what's left of the term, or written off if the commitment truly has no further use. Flag any reserved commitment that's underused well before the term renews, so finance isn't the last to know.
Renegotiating before the term renews
Reserved commitments auto-renew far more often than teams expect, and a commitment sized for last year's workload can quietly overshoot this year's actual usage. Put the renewal date on a calendar the same way you would a lease, and revisit actual usage against the commitment at least one full billing cycle before it renews, so there's time to right-size it instead of auto-renewing into another year of an outdated bet.
Working through a blended rate example before committing
Say your training workload runs steady in the background most weeks, with bursts around a launch that roughly double demand for a short stretch. Sizing the reserved floor purely on the steady weeks leaves the burst covered only by full on-demand pricing, since spot capacity isn't reserved by definition and can simply not be there when demand spikes hardest.
Model both pieces separately before signing anything: price the steady floor at the reserved rate, and price the burst at a blended assumption that mixes spot with a fallback to on-demand for whatever spot can't cover in time. If the burst is frequent enough to plan around, a smaller secondary reservation sized to the burst, rather than the full peak, usually beats paying full on-demand rates every time it happens.
The mistake to watch for is sizing the reserved commitment to the peak instead of the floor, on the logic that it covers everything. That looks safer on paper and is actually the more expensive path, since it pays the higher committed rate for capacity that mostly sits idle between bursts, which is the exact idle-capacity problem a reserved commitment is supposed to avoid in the first place.
What Good Looks Like
Good looks like a GPU bill you can split by team and by commitment type without asking engineering to reconstruct anything after the fact.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Frequently Asked Questions
How much of our GPU workload should be reserved versus spot?
Size the reserved floor to your lowest observed steady state usage over the last two or three months, and let spot cover everything above that. If your workload is genuinely variable with no clear floor, lean more heavily on spot until a pattern emerges.
How do we allocate a blended GPU bill to specific product teams?
Tag every job or endpoint with a cost center at launch time, inside your infrastructure provisioning, rather than trying to reconstruct ownership from the bill later. Cloud provider cost allocation tags exist for exactly this and are far cheaper than manual reconciliation.
Should a one-year GPU commitment be capitalized?
A prepaid commitment is usually recorded as a prepaid asset and amortized over the commitment term rather than expensed all at once, but the exact treatment depends on how the contract is structured. Check the specific terms with your accountant before you book it.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Modeling the Power Bill Behind Your GPU Cluster
How PUE, demand charges, and utility contract structure turn a GPU cluster's power draw into an actual operating cost, and how to model it correctly.
Cutting Cloud Egress Fees Without Losing Multi-Cloud Visibility
Where egress charges actually come from, four practical safeguards to cut them, and how to give finance visibility into data transfer spend across clouds.
Edge vs Cloud AI Inference: When On-Device Actually Pays Off
How to find your own crossover point between on-device AI inference and a cloud API, once you count hardware, model limits, and update infrastructure.
Building a Cloud Tagging Taxonomy That Actually Sticks
Why a tagging policy in a wiki page decays within a quarter, and how to enforce a small, mandatory tag set in your deploy pipeline instead.
Active-Active vs Active-Passive: What Disaster Recovery Costs
The real steady-state cost gap between active-active and active-passive disaster recovery, and how to size the decision around your actual downtime cost.
What AWS and Azure Marketplace Listings Actually Cost You
Learn what AWS and Azure marketplace listings cost beyond the fee: listing work, co-sell rules, payout timing, reconciliation and sales commission effects.