Human Labeling or Synthetic Data: Comparing the Real Cost per Usable Example
A vendor's per-label price is the least useful number in a data procurement decision, because it ignores quality variance, rework, and the review cycle needed to turn a raw label into something you'd actually train on. Synthetic data avoids the per-label fee entirely but introduces its own review burden to catch generation artifacts and bias before they end up baked into a model.
The comparison that actually matters is cost per usable example, after review and rework, not the sticker price of either approach.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Why does the quoted per-label price understate real cost?
A labeling vendor's quote covers the labeling task itself, not your internal cost to review a sample for quality, catch systematic errors, and send batches back for rework. Depending on the task's complexity and the vendor's quality track record, that review and rework cost can be a meaningful addition on top of the quoted rate, and it's almost never included in the number used to compare vendors or compare against a synthetic alternative.
Ask any vendor you're evaluating for their actual rework rate on comparable past projects, not just their quoted price, and treat a vendor that can't or won't share that number as a real data point in itself about how the comparison is likely to hold up in practice.
Why synthetic data isn't free either
Generating synthetic examples avoids a per-label fee, but replaces it with compute cost for generation and, critically, a review process to catch cases where the synthetic data doesn't actually reflect the distribution or edge cases your real-world data would. Skipping that review to save cost is how a model ends up trained on synthetic examples that quietly don't match production reality, which shows up much later as a model quality problem, not a line item you can trace back to the shortcut.
Budget the review step into the synthetic data cost estimate from the start, the same way you'd budget rework into a labeling vendor estimate, rather than treating synthetic generation's low direct cost as the full picture of what it actually takes to get to a usable, trustworthy dataset.
How do you compare labeling and synthetic data fairly?
Define "usable" concretely before comparing anything: passes your quality review, matches your labeling schema correctly, and reflects the distribution of real cases you need covered. Run a fixed-size pilot through both human labeling and synthetic generation against that same definition, and calculate total cost, including your internal review time, divided by the count that actually passed. That number, not the quoted price per unit, is what should drive the procurement decision.
Keep the pilot's sample size large enough that the rejection rate is a real signal rather than noise from a handful of unlucky examples, since a very small pilot can make either approach look better or worse than it would at real production volume.
To compare both approaches on a usable-example basis, work through these steps:
- Define a usable example in writing: it passes your quality review, matches your labeling schema, and reflects the mix of real cases you need covered.
- Run the same fixed-size pilot through human labeling and through synthetic generation, holding both to that one definition of usable.
- Add your internal review and rework time to each approach's quoted or compute cost to get an all-in total.
- Divide each all-in total by the count of examples that actually passed review to get the cost per usable example.
- Ask each vendor for its rework rate on comparable past projects and compare it against what your pilot showed.
Where the two approaches typically diverge
Human labeling tends to handle rare, nuanced, or judgment-heavy cases better, since a trained human can apply context a generation process can't easily replicate. Synthetic data tends to win decisively on cost for high-volume, well-understood, common cases where the pattern is easy to generate correctly at scale. Most real procurement decisions aren't a clean either-or once you look at the actual case mix you need covered.
Negotiating with a labeling vendor on quality, not just price
Once you have a real rework rate from a pilot, that number becomes useful in vendor negotiation, since a vendor whose quoted price is lower but whose rework rate is meaningfully higher may cost more per usable example than a pricier vendor with a cleaner track record. Bring the rework data to the negotiation rather than negotiating on the quoted rate alone, and ask vendors to commit to a quality threshold, not just a price.
A worked example that shows why rework rate matters more than price
Say Vendor A quotes a lower per-label rate but your pilot shows a meaningful share of its batches need rework before they clear your quality bar, while Vendor B quotes a higher per-label rate with a much lower rework rate on the same pilot. Divide each vendor's all-in cost, the quoted labeling cost plus your internal review and rework time, by the count of examples that actually passed on the first pass or after rework, and the ranking between the two vendors can flip once you look at it this way, even though Vendor A still wins on the number printed on the quote.
The mistake procurement teams make most often is comparing quoted rates across vendors and picking the lowest one, then discovering the real cost gap only after volume ramps up and the review team is buried in rework from the cheaper vendor. Run the pilot before committing to volume specifically to catch this, since a rework problem that looks manageable at pilot scale can become a real bottleneck once it's multiplied across a full production batch.
Share the worked comparison with whoever approves the vendor budget, not just the win-lose conclusion, so the next procurement decision starts from the same rigor rather than reverting to comparing quoted rates once this particular pilot is forgotten.
What Good Looks Like
Good looks like a cost per usable example figure, built from a real pilot, driving the procurement decision rather than a vendor's quoted per-unit rate.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Frequently Asked Questions
How do we estimate cost per usable example before running a full procurement?
Run a small pilot batch through your actual review process for both approaches before committing to a larger volume. The pilot's rework and rejection rate, applied to the quoted or estimated per-unit cost, gives a much more realistic cost per usable example than the vendor's raw quote alone.
Is a blend of human labeling and synthetic data usually the right answer?
Often yes, especially using synthetic data to cover common, easy-to-generate cases at low cost while reserving human labeling for rare or nuanced edge cases where synthetic generation is less reliable. This tends to lower blended cost per usable example compared to either approach used exclusively.
Should we track cost per usable example ongoing or just at procurement time?
Ongoing, if data collection is a recurring part of your model development process. Vendor quality and your own review standards both drift over time, and a cost per usable example that looked good at the initial procurement decision can quietly get worse without a periodic recheck.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Where RAG Pipelines Actually Rack Up Data Transfer Costs
Cross-region egress, repeated re-embedding, and vector storage bloat are the quiet costs behind a RAG pipeline, and a checklist for catching each one.
Startup Data Room Checklist: What to Prepare Before Diligence
Build a fundraising data room investors can review quickly: folder structure, what goes in each folder, access controls and the gaps that slow closings.
Automating Storage Tiers So Cold Data Stops Costing Hot Prices
How automated lifecycle rules move aging data to cheaper storage tiers on their own, and the retrieval cost tradeoffs worth checking before you set them.
What AWS and Azure Marketplace Listings Actually Cost You
Learn what AWS and Azure marketplace listings cost beyond the fee: listing work, co-sell rules, payout timing, reconciliation and sales commission effects.
Is Your Copilot Seat Actually Paying for Itself?
A practical way to measure whether AI coding tool seats are worth their cost, beyond adoption rate, using output and review metrics you already track.
Calculating a Real Cost-Per-Transaction Number
Why total infrastructure spend hides whether growth is healthy, and how to build a cost-per-transaction number that survives a shifting mix of usage.