LLM product unit economics: an evidence-first pricing worksheet
Advanced13 min readAI for Business

LLM product unit economics: an evidence-first pricing worksheet

Model usage distribution, contribution margin, failure handling, support, and retention before choosing a price. This worksheet replaces unsupported market ranges with auditable inputs.

What you should be able to do

Price from measured customer value and the full cost-to-serve distribution—not average tokens alone. Treat defensibility claims as hypotheses and test retention, switching behavior, and willingness to pay.

Saved only in this browser.
In this article

LLM products can carry material variable cost and third-party platform dependency. Whether those costs make the business weaker depends on actual usage distribution, pricing, support load, retention, and customer value.

This article teaches how to build an auditable unit-economics model and test possible sources of defensibility. It does not report an industry survey or predict which companies will survive.

The FinOps Foundation defines unit economics as connecting technology cost to a business-value metric. Use that as a comparable operating framework, then have finance define the actual unit, cost allocation, and margin treatment for this product.

What’s different about LLM products

A few characteristics that distinguish LLM-powered products from traditional SaaS:

Usage can create material marginal cost. Model calls, tools, retrieval, storage, review, and support can rise with use. The marginal cost is not necessarily equal for the first and ten-thousandth call because caching, batching, discounts, routing, and capacity utilization change it; calculate the distribution from traces and invoices.

Margins depend on the full cost-to-serve distribution. Do not import an uncited industry range. Define gross margin and contribution margin with the finance reviewer, then calculate both from the company’s ledger and usage records.

Foundation-model access may be shared. A public model alone may be easy for competitors to obtain, but implementation quality, contracts, data rights, distribution, operations, and customer trust can still differ.

Foundation models change. Versions, deprecations, prices, quotas, and terms can change on provider schedules. Record each dependency, notice path, fallback, migration test, and commercial trigger.

Features can be copied. Test whether implementation quality, rights-cleared data, distribution, contracts, integration, operations, or trust create observed customer value; do not assume a prompt, tune, or workflow is either unique or trivial to reproduce.

Customer expectations and vendor features change. Re-test willingness to pay and competitive alternatives on a defined schedule.

Competitive overlap. Model, cloud, SaaS, and specialist vendors may add adjacent features. Maintain a dated competitor map and test customer alternatives rather than inferring pressure from company size or partnerships.

Treat these as risks to model, not facts about every LLM product.

The cost structure

Build the product’s cost structure from its ledger and traces. Start with cost pools, then have finance classify their behavior by activity base, contract, capacity commitment, and decision horizon; the same cost can be fixed for one decision and variable or step-fixed for another.

Potential metered or volume-sensitive pools:

  • Inference (LLM API or hosting).
  • Embedding (for RAG).
  • Vector DB or storage.
  • Other infrastructure that scales with usage.

Potential period, committed-capacity, or step-fixed pools:

  • Salaries and support staffing, depending on hiring and service-volume decisions.
  • Office and operational commitments.
  • Software licenses, which may be flat, per-seat, tiered, or usage-based.
  • Hosting baseline or reserved capacity, including step changes.
  • Marketing and sales, separating committed spend from commissions or activity-driven costs.

Calculate contribution per account as recognized revenue minus the agreed variable and attributable costs. Include model/tool usage, retrieval and storage, human review, payment fees, support, refunds or credits, and any other cost finance classifies as variable or attributable.

Model the distribution rather than one invented average. At minimum, calculate cost-to-serve by account and by usage percentile, then stress-test current provider prices, retries, longer context, model fallback, support tickets, and abuse. Re-check prices from the providers’ official pages before every decision cycle.

Pricing models

Pricing requires explicit treatment of usage-sensitive cost, customer value, and billing operations. Candidate models include:

Flat per-user. Predictable to quote, but usage variance can create cross-subsidy or loss-making cohorts.

Per-seat with usage caps. Each seat includes N actions/calls/tokens per month. Heavy users pay overage or hit limits.

Pure usage-based. Pay per call, token, action, or other unit. It may align revenue with one cost driver but does not guarantee predictable margin; support, retries, discounts, minimum commitments, and value can differ by account.

Tiered. Packages can separate capability, service, limits, or support. Validate whether the tiers are understandable, contractually operable, and economically distinct.

Hybrid. Per-seat baseline plus usage-based allowances or overage. It may balance predictability and cost alignment, but only cohort economics and customer research can establish that.

Outcome-based. Pay per defined outcome. It can connect price to value but creates attribution, quality, dispute, fraud, timing, and accounting questions; this article has no market dataset proving an adoption trend.

Each has trade-offs. The “right” model depends on:

  • Predictability needs (yours and the customer’s).
  • Variance in usage.
  • Margin structure.
  • Competitive landscape.

Do not assert a dominant winning model without market evidence. Test candidate pricing with customer research, legal review, billing feasibility, and cohort economics.

The unit economics question

A useful frame: what’s the unit you charge by, and what does it cost?

For a chat assistant, a candidate unit may be a conversation, resolved task, seat, or usage allowance. Measure the model/tool cost, human review and support, failure/retry cost, and revenue attributable to that unit.

The exercise: measure cost per unit, test revenue and packaging candidates, and evaluate the finance-approved margin target across the usage distribution.

High-usage accounts may be more costly and may also retain, expand, or create more value. Analyze contribution and retention by cohort before changing limits.

Margin protection

A few tactics for protecting margins:

1. Tier the model by feature

Cheap-to-run features available on basic tier. Expensive features (reasoning models, long context, large outputs) on premium.

Measure cost by feature and route. If a premium capability has a materially higher cost-to-serve or customer value, test whether a separate allowance or tier is understandable and commercially viable.

2. Cost optimization (covered in Cost-optimizing inference)

Caching, routing, and output control are candidates. Use the trace-based method in the linked article; do not budget a savings percentage before measuring.

3. Usage transparency

Show users their usage. Implicit guidance for them to optimize their own behavior.

Usage visibility can help customers understand limits and charges, but it can also confuse or discourage use. Test comprehension, accessibility, behavior, support load, and conversion rather than assuming users self-throttle or upgrade.

4. Smart caching

User- or organization-scoped caching may reduce repeated work when freshness, privacy, authorization, invalidation, and hit-rate tests support it. Measure net cost and user outcomes; do not share personalized cache entries across scopes.

5. Hybrid hosted/self-hosted

Self-hosted or BYO-cloud variants change who operates and pays for infrastructure. They can also increase support, security, release, and compatibility cost; model the contract as a whole.

6. Outcome-based for high-value

Some workflows have measurable outcomes. Outcome pricing changes attribution, dispute, fraud, timing, and revenue-recognition questions; qualified financial and legal review is essential.

The moat question

The harder question. What makes your product defensible?

A list of hypotheses to test against customer behavior:

Candidate sources of defensibility

Lawfully controlled data. Approved data rights, quality, and feedback may improve the product. Test whether this changes outcomes or switching behavior; do not trap customers or obscure export/deletion rights.

Distribution. Existing access to a target segment may lower acquisition friction. Measure conversion and retention rather than assuming reach creates defensibility.

Trust. Measure whether security evidence, domain review, reliability, support, and responsible incident handling affect acquisition, renewal, or expansion. A regulated-industry label does not establish trust.

Integration. Integrations may reduce user friction and create operating value. Measure it without designing unfair lock-in, and preserve export and offboarding.

Workflow specialization. A domain-specific implementation may outperform a generic alternative, but only task results, adoption, reviewer evidence, and switching research establish an advantage.

Network effects. Define the mechanism by which another participant increases value, obtain the necessary data rights, and measure the effect. A community or shared dataset is not automatically a network effect.

Brand and switching evidence. Use win/loss research, renewal behavior, migration effort, customer-controlled export, and trust measures. Do not speculate about a buyer’s career or treat customer difficulty leaving as the goal.

Vertical integration. Owned models, inference, or data pipelines may improve control or economics, but can also increase capital and operations burden. Measure the claimed advantage.

Pseudo-moats

A specific prompt or workflow. Treat copyability and customer build capability as research questions, not assumptions.

A specific public model choice. Access alone is rarely exclusive, although contracts, regions, tuning, serving, and operations can still differ.

A clever UI. UI alone is not evidence of defensibility; neither is an invented copying timeline.

Speed of iteration. It may create a temporary advantage, but durability requires evidence from customer value, retention, operations, or another reinforcing mechanism.

Marketing or branding alone. Brand can matter, but its defensibility must be demonstrated through acquisition, retention, trust, or pricing evidence.

The weaker hypotheses are easier to copy or may decay quickly. Do not connect them to a mortality rate without a cited dataset.

The strategic patterns

Patterns that may contribute to defensibility:

Pattern 1: Workflow + AI, not “AI tool”

Rather than “AI that does X,” build a workflow that incorporates AI as part of a larger system.

Example: not “AI summarizer for legal contracts,” but “contract management platform with built-in AI.”

Test whether the surrounding workflow improves activation, retention, willingness to pay, or switching behavior. A larger feature set does not automatically create a moat.

Pattern 2: Customer-data flywheel

Each customer’s usage produces data that improves their experience (and possibly others’). Switching means losing the accumulated personalization.

Example hypothesis: an AI sales assistant that uses approved account context and confirmed preferences could reduce repeated setup. Test portability, customer control, and whether the benefit persists.

Build only with explicit rights, customer controls, isolation, correction, deletion, and a measured improvement loop. Accumulating data without those controls is liability, not a flywheel.

Pattern 3: Deep vertical specialization

Pick a vertical. Build for it deeply. Healthcare, legal, finance, real estate.

Vertical knowledge may improve workflow fit. It also raises domain-review, liability, data, and compliance requirements. Measure advantage against both general and specialized competitors.

Pattern 4: Embed in existing workflow

Embedding into an approved system where users already work is one candidate for reducing context switching; it is not always preferable to a separate product boundary.

Embedding can reduce workflow friction and increase integration dependency. Measure usage and retention, and preserve fair export and offboarding behavior.

Pattern 5: AI-native operations

Some businesses are AI-native end-to-end — they don’t sell AI to humans; they use AI to deliver a service. AI tutoring, AI customer service as outcome, AI content as commodity.

The product may be the service outcome while AI remains an operational component. Compare quality, cost, reliability, and customer preference with alternative delivery models.

Pattern 6: Multiple reinforcing hypotheses

Multiple advantages may reinforce each other. Establish each link with evidence rather than labeling the combination successful in advance.

Evidence records instead of fictional case studies

Do not fabricate an anonymized startup or financial outcome. For each pricing experiment, keep a decision record with:

  • hypothesis and customer segment,
  • price, allowance, overage, cancellation, and refund terms,
  • sample size and experiment dates,
  • activation, usage distribution, conversion, retention, expansion, support, and churn,
  • cost-to-serve distribution and contribution definition approved by finance,
  • qualitative customer research and known selection bias,
  • legal, tax, billing, and consumer-protection review,
  • decision, confidence, owner, and next review date.

For defensibility, record evidence such as renewal behavior, time-to-value, integration depth, approved data rights, switching interviews, sales-cycle effects, and competitor win/loss reasons. A plausible story is not evidence of a moat.

A person compares an unbranded product prototype with blank test cards
AI-generated illustration of validating a product prototype and recording feedback before release.

What goes wrong

Illustrative risk scenarios to test:

Failure 1: Margin compression. Started with healthy margins; competitors and price pressure squeezed them. Now revenue grows but profit doesn’t.

Failure 2: Provider substitution. A foundation-model or platform provider launches a sufficient competing feature, and measured differentiation declines.

Failure 3: High-usage cohort economics. Some accounts disproportionately drive cost. Test pricing, limits, workflow design, model routing, support, and value before deciding whether or how terms should change.

Failure 4: Quality regression. Foundation model update changed behavior. Your tuned prompts broke. Customer trust dropped. Recovery is slow.

Failure 5: Customer substitution. A lower-cost platform or competitor becomes sufficient for the workflow, and retention declines.

Failure 6: Unmanaged substitution risk. A platform or competitor adds a sufficient alternative, but the product has no dated competitor review, customer evidence, or migration plan.

Failure 7: Scaling without margin discipline. Growth funded growth; margins ignored. Eventually money runs out, no path to profitability.

Failure 8: Foundational dependency risk. API price increase, deprecation, or outage from your provider. Your business is disrupted by something you don’t control.

Pricing hypotheses to test

Per-action with caps. Test whether the action is unambiguous, auditable, valuable, resistant to gaming, and aligned with cost and customer expectations.

Bring your own API key. This can move a provider charge to the customer, but does not remove support, security, integration, failure, or accounting implications. Verify provider terms and tenant isolation.

Negotiated pricing. Test whether expected volume, service, support, risk allocation, and discounts produce an approved contribution margin under adverse usage.

Free trial or tier. Measure activation, conversion, abuse, support, infrastructure cost, retention, and cannibalization. Free usage does not promise paid conversion.

Term commitments. Model billing, revenue recognition, discounts, minimum usage, service obligations, cancellation, collection, renewal, and forecast error with finance and legal reviewers.

A framework for the pricing decision

Steps:

  1. Model your cost distribution. What does it cost to serve each account and meaningful usage percentile, including failures and support?

  2. Choose your unit. What are you charging for? Seats, actions, outcomes, tokens?

  3. Test price candidates. Measure willingness to pay, conversion, retention, value evidence, cost, and competitive alternatives by segment.

  4. Test packaging. Use only tiers and limits that customers understand and systems can enforce and bill accurately.

  5. Set limits. Where do heavy users start hurting margins? Enforce caps or charge overage.

  6. Plan for change. Foundation-model prices and features can rise, fall, or change shape. Define how vendor changes trigger pricing and product review.

  7. Measure on a defined cadence. Have finance define customer lifetime value, contribution margin, churn, and expansion before using them in decisions.

How to research the market claim

Avoid broad statements such as “wrappers fail,” “vertical products thrive,” or “platforms dominate” unless a named dataset defines the population, period, and outcome. Build the product’s evidence from:

  • current competitor capabilities and prices captured with dates,
  • customer win/loss interviews,
  • cohort retention and expansion,
  • vendor roadmap and deprecation monitoring,
  • switching and integration costs observed in customer research,
  • cited industry datasets whose methodology is suitable for the claim.

Separate a strategic hypothesis from a measured market finding in every review.

What to optimize for

For founders, a hierarchy:

  1. Build something that does real work. Not “AI for X.” A measurable outcome customers value.

  2. Test defensibility hypotheses deliberately. Customer value, lawful data advantage, vertical depth, distribution, integration, and workflow embedding all require evidence.

  3. Validate unit economics. Stress-test approved margin definitions and price candidates across real cohorts and forecast ranges.

  4. Manage foundation dependencies. Diversification and portability have costs; choose fallback and migration controls from impact and tested feasibility.

  5. Scale operational evidence. Observability, evaluations, security, quality, incident handling, and recovery make reliability and cost claims auditable.

  6. Earn durable relationships. Trust, reliable integrations, value, and responsible offboarding matter. Do not treat customer lock-in as the goal.

The underlying vendor landscape changes quickly, so keep review dates and exit plans explicit.

Build with the moat in mind

Shipping an LLM product requires explicit variable-cost and dependency management. Whether its margins or competition are better or worse than another SaaS category must be established from comparable data.

Possible sources of defensibility include lawful data advantages, vertical depth, distribution, integration, trust, and workflow fit. Treat each as a testable hypothesis.

The economics need rigor. Attribute variable and avoidable costs, test caching/routing/packaging changes against quality and policy, and handle high-cost cohorts through reviewed product and commercial decisions rather than a generic cap rule.

The output of this worksheet is an auditable pricing and defensibility decision with named assumptions, evidence, owners, and review dates. Qualified financial and legal reviewers must approve the definitions and customer terms before production use or external financial claims.

Read next

Continue through the same learning path with the next practical articles.