Open Source vs Proprietary LLMs: What Each Choice Costs Your Margin
"Open source is free" is the most expensive sentence in this decision. A proprietary API charges per token and nothing else. An open source model charges you in GPU hosting, the engineering time to run and tune it, and the ongoing burden of keeping it patched and performant, none of which shows up as a line item until the team is already committed.
The right call depends on your volume and your margin structure, not on which option sounds more strategic in a board deck.
What the proprietary path actually costs
A proprietary API's cost is close to fully variable: you pay per token, the vendor handles scaling and uptime, and your engineering cost is limited to integration and prompt work. The tradeoff is that the per-unit rate is fixed by someone else, and at high volume that rate can outgrow what a self-hosted alternative would cost. It's also the option with the least operational risk, since you're not the one paging an on-call engineer at 2am for a model server.
What the open source path actually costs
Open weights remove the per-token licensing fee, but you now own GPU hosting, model serving infrastructure, monitoring, and periodically re-tuning as your use case shifts. Add up hosting cost per hour, the engineering hours to stand up and maintain serving infrastructure, and the fully loaded cost of whoever owns on-call for it. Below a certain volume, that fixed cost is more expensive per request than a proprietary API would have been.
For example, a company with one narrow classification feature and a small platform team might assume self-hosting will cut its bill. Once it counts the on-call rotation, the time spent patching and re-tuning the model, and the GPUs that sit idle overnight, the cost per request can land above what the proprietary API would have charged at that volume. A simple decision rule follows from this: self-host only when volume is high and consistent enough to keep the hardware busy, or when a compliance requirement leaves no alternative. If neither is true, stay on the proprietary API and revisit the comparison as volume grows.
Doing the margin math
Take your current or projected monthly token volume, price it against the proprietary rate, and separately estimate the fully loaded monthly cost of self-hosting at that volume, including engineering time. Where the two lines cross is your breakeven volume. Below it, proprietary wins on pure cost. Above it, self-hosting starts pulling gross margin the other direction, assuming your team can actually operate the infrastructure reliably.
Run the comparison in this order:
- Estimate your current or projected monthly token volume and price it at the proprietary rate.
- Estimate the fully loaded monthly cost of self-hosting at that volume, including GPU hosting, serving infrastructure, monitoring and the engineer who owns on-call.
- Find the volume where the two cost lines cross, which is your breakeven point.
- Add switching cost, meaning the engineering time to move providers or bring hosting in-house, as its own line item.
- Set a calendar reminder to rerun the comparison as proprietary rates, open model quality and your own volume change.
Where the decision usually reverses
Two situations flip the answer regardless of volume. First, data residency or contractual requirements that rule out sending data to a third party API, which makes self-hosting a compliance decision rather than a cost decision. Second, a use case narrow enough that a smaller fine-tuned open model matches a much larger proprietary model's quality at a fraction of the inference cost, which changes the whole comparison in the open source direction.
Building the switch cost into your model
Whichever way you go, don't treat it as permanent. Model your switching cost, meaning the engineering time to move providers or bring hosting in-house, as its own line item. A margin comparison that ignores switching cost will consistently undercount the value of staying on a proprietary API a little longer, since flexibility has a real, if hard to quantify, price. Software companies generally carry meaningfully higher gross margin than services businesses, and this decision is one of the places that margin gets protected or eroded.
Revisit the comparison on a schedule rather than once. Proprietary rates change, open model quality improves, and your own volume grows, so a decision that was clearly right a year ago can quietly stop being right without anyone rerunning the numbers. Put a calendar reminder on it the same way you would a lease renewal.
What this looks like on the income statement
Whichever path you pick shows up in a different place. Proprietary API spend is a clean, fully variable cost of goods sold line that scales directly with usage, which investors and boards generally find easy to read. Self-hosting splits your cost across infrastructure spend and the fully loaded cost of the engineers maintaining it, some of which may not even sit in cost of goods sold depending on how your team is structured. Decide up front how you'll classify self-hosted infrastructure and engineering time so your gross margin is comparable quarter to quarter regardless of which path you're on.
What Good Looks Like
Good looks like a documented breakeven volume comparing both paths at your actual usage, refreshed whenever provider pricing changes.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
What volume typically makes self-hosting worth it?
There's no universal number, since it depends on your specific proprietary rate and hosting cost, but most teams find the crossover happens well above what an early stage product sees. Build your own breakeven model rather than borrowing someone else's threshold; the engineering time to maintain serving infrastructure is the cost most teams underestimate.
Does using an open source model reduce our compliance burden?
Not automatically. Self-hosting can help with data residency requirements since data never leaves your infrastructure, but you take on the compliance burden of securing that infrastructure yourself. Whether that's a net simplification depends on your existing security posture.
Can we run a hybrid setup with both model types?
Yes, and many teams do: proprietary for general capability and low volume features, self-hosted or fine-tuned open models for the one or two high volume, narrow tasks where the margin math clearly favors it. That split usually beats an all-or-nothing choice.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
The Margin Math Behind Open-Core Software
Why the paid tier of an open-core business has to cover more than its own delivery cost, and how licensing choices shape that margin for years afterward.
How Prompt Caching Actually Cuts Your LLM API Bill
How cache hit rate turns into real savings on your LLM API bill, a worked example, and what quietly breaks cache performance in production.
When You Can Capitalize LLM Fine-Tuning Costs Under ASC 350-40
How ASC 350-40's three development stages apply to LLM fine-tuning and RAG pipeline work, so you know which costs to expense and which to capitalize.
What Happens to Your Margin When Your Model Provider Raises Prices
Why AI wrapper products are exposed to upstream price hikes, how to see the exposure coming, and the contract and product levers that protect margin.
Capturing the Batch API Discount Without Hurting Your Product
How to find the AI calls that can tolerate a delay, move them to a discounted batch endpoint, and recover real margin without touching real-time features.
Where Observability Budgets Actually Leak, and How to Contain Them
The specific places logging, metrics and tracing spend leaks in a growing engineering org, and the containment moves that work without cutting visibility.