Cost-optimizing inference: prompt caching, routing, and output control
Advanced12 min readAI for Business

Cost-optimizing inference: prompt caching, routing, and output control

Build a trace-based inference cost model, optimize the largest measured contributors, and prove that each change preserves task quality.

What you should be able to do

There is no universal savings percentage. Measure tokens, cache hits, retries, tools, latency, and quality by workflow; then optimize the largest cost line and re-evaluate.

Saved only in this browser.
In this article

Inference cost is a workload property: model and region, uncached and cached input, generated and reasoning tokens, tools, retries, concurrency, storage, and human review all matter. Start from billed usage and traces rather than an industry savings claim.

Current prices and cache rules change frequently. Re-check the official OpenAI pricing, Anthropic pricing, and Gemini pricing pages before using any cost model.

This article covers the techniques, the numbers, and the production discipline. We assume you’ve already done basic model routing (covered in Multi-model orchestration); we’re going deeper.

The cost stack

LLM costs come from:

  • Input tokens. What you send to the model. Includes system prompt, context, user query.
  • Output tokens. What the model returns. The input/output price ratio varies by provider, model, batch mode, and cache state.
  • Reasoning or hidden-compute charges. Provider reporting and billing treatment varies; use the billed usage fields and current price sheet rather than assuming parity with visible output.
  • Tool definitions and tool use. Schemas may add input tokens, while hosted tools or external services may have separate charges.
  • Retries and failed work. Some failed or abandoned calls incur usage; classify failure points from provider billing and traces.

Optimization works at each layer.

Technique 1: Prompt caching

Caching can be a major lever when requests share an eligible prefix and the workload produces a high hit rate.

Read each provider’s current cache documentation for minimum prefix size, write/read charges, expiry, routing constraints, and observability.

At a high level, an eligible repeated prefix may be reused within a provider-defined window. The exact semantics are not portable across providers.

Practical implementation:

Structure your prompts so static content comes first, dynamic content last:

[CACHED: 10K tokens]
- System prompt
- Tool descriptions
- User's static profile
- Knowledge base snippets unlikely to change per call

[NOT CACHED: 1K tokens]
- Conversation history (changes each turn)
- Current user query

Whether this prefix is eligible and how it is billed depends on the selected model and provider.

Savings equation:

daily cost = calls × (uncached input × input rate + cache writes × write rate + cache reads × read rate + output × output rate + tool charges)

Populate it with billed token counts and the current price sheet. Compare equivalent time windows and include cache misses.

Implementation discipline:

  • Identify static vs dynamic parts of prompts.
  • Place static parts first.
  • Use cache markers where the provider supports them (Anthropic) for explicit control.
  • Test cache hits — your observability should show cache hit rate. If it’s low, your prompt structure isn’t right.

Prioritize caching only after the trace shows repeated eligible input is a leading cost line.

Technique 2: Model routing

Covered in detail elsewhere. Briefly: different requests to different models based on complexity.

Build a labeled routing evaluation set, compare candidate models on quality and cost, and keep a fallback for uncertain or failed cases. The resulting routing mix is workload-specific.

Technique 3: Output length control

Output tokens may be a leading cost line. Confirm that in usage data before optimizing response length.

Strategies:

Explicit length instructions.

Respond in at most 100 words.

Models may still violate a prose length instruction. Enforce and evaluate the limit, and report the measured token and quality change rather than promising significant savings.

Structured output.

When the required response is short structured data, a strict schema can reduce irrelevant prose. It does not eliminate invalid output, oversized field values, retries, or truncation risk; validate every result.

Provider output-token cap.

Set the current API’s output cap from measured task needs and leave headroom for valid completion. Parameter names and semantics differ by API and model; an overly small cap can truncate structured output and cause more calls.

Format constraints.

“Bullet points only” or “single paragraph” produces shorter outputs than free-form.

Bullet over prose.

Bullets may reduce connective prose for some responses. Measure token counts; a verbose bullet list can be longer than a concise paragraph.

No preamble.

“Skip introductory phrases. Get straight to the answer.” Models often start with “Great question…” or “Let me explain…” — wasted tokens.

Measurement example:

For a summarization workflow, compare default and constrained outputs on the same documents. Report billed output tokens, factual coverage, readability, user follow-up rate, and any truncation. A lower token count is not a saving if users need another call.

Technique 4: Output sampling and early stop

For some use cases, you don’t need full LLM output — you need a decision or a classification.

Logprobs for classification.

# Use only a model and API operation whose current documentation supports
# log probabilities; support varies by model and endpoint.
response = openai.chat.completions.create(
    model=SMALL_NON_REASONING_MODEL,
    messages=[{"role": "user", "content": prompt}],
    logprobs=True,
    top_logprobs=5,
    max_tokens=1
)
# Read logprobs of first token to determine likely category

This asks a compatible model to emit a short label. Confirm that the selected API supports log probabilities, that labels map cleanly to tokens, and that classification quality meets the target.

Constrained labels or logit bias.

For known-set outputs, prefer a documented enum/schema constraint where available. Logit bias influences token selection but is tokenizer- and model-specific.

response = openai.chat.completions.create(
    model=COMPATIBLE_MODEL,
    messages=[...],
    # If used, build logit_bias from every tokenization variant you intend
    # to accept; do not assume a class label is exactly one token.
    logit_bias=VALIDATED_TOKEN_BIAS,
    max_tokens=1,
)

Logit bias influences token selection; it does not enforce a valid class or prove classification reliability. Validate and reject unexpected output.

Technique 5: Batching

When you’re processing many items, batch them.

Async batching at the API level.

Some providers support asynchronous or batch products, sometimes with different prices, completion windows, limits, and data-handling terms.

  • OpenAI and Anthropic both document asynchronous batch products. Verify the current discount, completion window, limits, and data-handling terms on their official pricing and batch pages.

If backlog work does not need an interactive response, compare current batch pricing and completion behavior with the synchronous path.

In-prompt batching.

Process multiple items in one LLM call when possible.

Instead of:

[10 separate calls, each classifying one ticket]

Do:

[1 call, classifying 10 tickets in one prompt]

The single call has more item input but may reuse one set of fixed prompt overhead. It can reduce total tokens in some workloads; separators, longer outputs, retries, and batch-level failure can erase that saving.

Caveat: batching can change quality, ordering, truncation, and failure isolation. Sweep batch sizes on representative inputs instead of adopting a universal range.

Technique 6: Smaller models for narrow tasks

Beyond standard routing — consider whether a task really needs a big model.

Classification: evaluate a smaller tier against the current production model on labeled examples, including rare classes and abstention. Use current provider prices in the cost comparison.

Extraction: Compare smaller, mid-tier, deterministic, and hybrid extractors on field-level accuracy, exception handling, latency, and cost. Escalate cases using tested rules.

Translation: Evaluate specialized translation systems and LLM tiers on the actual language pairs, terminology, formatting, safety, and human-review requirements. Do not infer coverage from aggregate benchmarks.

Embedding: Use embedding-specialized models, not general-purpose LLMs for embedding.

The pattern: identify your “simple, narrow” workloads. Route them to the smallest model that does the job adequately. Save flagship for the complex, judgment-heavy work.

Technique 7: Fine-tuned small models

For very high-volume narrow tasks, fine-tune a small model.

For a high-volume classification workload, compare a prompted small model, a fine-tuned model, deterministic rules, and a hybrid. Include training and evaluation data, serving, idle capacity, monitoring, retraining, and engineering cost. Fine-tuning is economical only if the measured quality/cost curve justifies it.

We covered this in Fine-tuning in 2026. The principle: when scale and narrowness align, fine-tuning is a cost lever.

Technique 8: Pre-filtering

For multi-step LLM workflows, cheap filtering catches obvious cases before expensive processing.

Example: customer support classification + response.

Cheap pre-filter:

  • “Is this an actual support question or spam/noise?” (1-token classification on a small model.)
  • “Is this a known FAQ?” (Embedding search; cheap.)

Only requests passing the filter reach the expensive response generation.

Measure what share of traffic the pre-filter can resolve at the required precision. False positives may suppress valid requests, so cost savings must be evaluated alongside quality and escalation impact.

Technique 9: Caching beyond prompt caching

Beyond the model provider’s prompt caching, application-level caching:

Response caching. For an approved scope and complete set of response-affecting inputs, reuse a prior response when its freshness and semantics remain valid. Non-determinism means this is a product decision, not an identity law.

import hashlib
import json

def stable_sha256(value):
    payload = json.dumps(value, sort_keys=True, separators=(",", ":"))
    return "llm:" + hashlib.sha256(payload.encode("utf-8")).hexdigest()

def cached_call(scope, prompt_version, prompt, model, params, ttl=3600):
    # Canonicalize all response-affecting inputs and include tenant/user scope
    # where a shared answer is not explicitly safe. Use a stable cryptographic
    # digest rather than the process-randomized built-in hash().
    cache_key = stable_sha256({
        "scope": scope,
        "prompt_version": prompt_version,
        "prompt": prompt,
        "model": model,
        "params": params,
    })
    cached = redis.get(cache_key)
    if cached:
        return json.loads(cached.decode("utf-8"))
    response = call_llm(prompt, model, params)
    redis.set(cache_key, json.dumps(response), ex=ttl)
    return response

For eligible deterministic-enough queries, this can avoid repeat calls after a populated hit. Define freshness and invalidation, isolate scopes, prevent cache stampedes, and do not cache sensitive or personalized output without approval.

Embedding caching. Computed embeddings cached.

Retrieval result caching. Search results for a query cached for short periods.

Tool result caching. Tool call results cached if the underlying data doesn’t change often.

Caching levels stack. At each layer, you save calls.

Technique 10: Speculative execution (latency trade-off, not a cost saving)

For latency-sensitive flows where you can predict next steps, speculatively pre-call.

Example: customer support agent. You know the next step is usually “summarize the issue” after the customer describes it. Start that summarization in parallel with showing acknowledgment to the user.

If the prediction is right, the response is ready when needed. If wrong, you wasted one call.

This deliberately spends work that may be discarded, so it can increase cost. Use it only when measured latency benefit justifies the waste, cancellation and side effects are controlled, and the speculative request cannot expose or mutate unauthorized data.

Technique 11: Provider comparison

Providers differ in price, capability, regions, quotas, data terms, reliability, and model implementation. Compare quality-equivalent candidates under the same workload and contract assumptions.

Open-weight models on managed inference providers.

Compare current provider prices only after the candidate model passes the same workload evaluation. Similar parameter size or marketing tier does not establish equivalent quality.

Same model on different providers.

Some open-weight models are offered by multiple providers. Benchmark the exact revision, quantization, serving configuration, and API behavior; the same model name does not guarantee identical output or performance.

Self-hosting at scale.

Self-hosting may become cheaper at a workload-specific utilization point. Model GPU hours, replicas, idle and peak capacity, networking, observability, upgrades, security, incident response, and engineering ownership.

Multi-provider routing adds integration, evaluation, security, procurement, observability, and failure-mode complexity. Adopt it only when the measured resilience or economic benefit exceeds that ownership cost.

Different task shapes are sorted beside a simple evaluation grid
AI-generated illustration of matching inference effort to representative task types.

Technique 12: Inference acceleration

For self-hosted: optimization of the inference layer itself.

vLLM, TGI, SGLang. Inference servers with different model coverage and optimization paths. Benchmark supported versions on the target hardware.

Quantization. Lower precision can reduce memory or improve throughput, with task- and method-dependent quality effects. Benchmark the exact artifact and serving configuration.

Flash Attention, paged attention. Architectural optimizations enabled in modern servers.

Continuous batching. Servers that batch in-flight requests for better GPU utilization.

For teams self-hosting at scale, this matters. For teams using APIs, the provider handles it.

Technique 13: Streaming

Streaming doesn’t reduce token count but improves UX, which matters for cost-effectiveness perception.

For long outputs, users see content appearing immediately. They can read along while generation completes. Feels much faster than waiting for full response.

For agents, stream approved progress events or status summaries. Do not expose private reasoning, secrets, unreviewed tool arguments, or cross-tenant data as “intermediate steps.”

Implementation: if the selected API and model support streaming, test it for user-facing flows. Streaming changes perceived latency, not necessarily total cost or task completion time.

Technique 14: Budget guards

Beyond optimization, enforce hard budgets to prevent runaway costs.

Per-request budget. Maximum tokens per request. Stop if exceeded.

Per-user budget. Daily or monthly cost cap per user. Throttle when approaching.

Per-feature budget. Each feature has an owned budget and a tested response when a workload-specific threshold is exceeded.

Global budget. Total daily/monthly limit. Pause non-essential work near limits.

These controls do not prove savings. They bound or redirect spend and can also reduce availability, so test alerting, throttling, downgrade, queueing, and circuit-break behavior for essential and non-essential workloads.

A cost-reduction experiment template

Do not present a composite as a customer result. For one production workflow, capture a baseline billing period and apply one change at a time:

The changes:

  1. Prompt caching. Record eligible prefix tokens, hit rate, cache writes/reads, latency, and billed cost.

  2. Model routing. Record route distribution, per-route quality, fallbacks, latency, and cost.

  3. Output length control. Record output length, completion quality, user re-prompts, and cost.

  4. Pre-filtering. Record precision, recall, escalations, suppressed valid requests, and avoided calls.

  5. Response caching for FAQ. Record semantic equivalence rules, freshness, invalidation, hit rate, and answer quality.

Report gross and net savings, evaluation results, engineering time, new operational cost, and uncertainty intervals where the data and method support them. Do not claim unchanged quality unless the evaluation design can detect meaningful regressions.

Common mistakes

Failure patterns to check in your own traces:

Mistake 1: No cost tracking. Team has no visibility into what each feature, user, or call costs. Optimization is impossible without measurement.

Mistake 2: Optimizing the wrong thing. Spent weeks reducing input tokens by 5% when output tokens were 80% of the bill. Measure first; optimize biggest contributors.

Mistake 3: Quality regressions. Cost cuts shipped without quality monitoring. Saved money, lost users. Always pair cost work with eval suites.

Mistake 4: Over-routing. Aggressive routing to small models for tasks they can’t really handle. False savings.

Mistake 5: Cache pollution. Cache filling with rare queries. Most cache entries used once. Cache misses dominate. Better caching strategy needed.

Mistake 6: Skipping batch evaluation. Real-time processing is used for work that could tolerate a provider’s current batch window and constraints.

Mistake 7: Over-engineering. Building elaborate cost optimization on top of features that aren’t profitable anyway. Sometimes the right answer is “kill the feature.”

Mistake 8: No budget guards. A single bug produces a runaway. Catastrophe rather than minor inconvenience.

Operating discipline

Recommended operating practices:

  • Treat cost as a metric, not an afterthought.
  • Assign an owner across engineering and finance responsibilities.
  • Review costs on a cadence suited to spend volatility and business risk.
  • Triage spikes against defined thresholds and runbooks.
  • Set budgets per feature; alert on threshold breaches.
  • Make trade-offs explicitly (cost vs quality vs latency).

Control gaps to look for:

  • No named cost owner.
  • Billing discovered only after the decision window has passed.
  • Alerts without a response owner or runbook.
  • No approved budget or forecast envelope.
  • Skip the trade-off conversation; optimize one dimension at a time.

These are governance choices to verify through ownership records, review attendance, alert response, and completed cost actions—not a claim about team culture.

Pricing and capability drift

A note on the broader trend.

Provider prices, model capabilities, batch products, cache rules, and hosted-tool charges change on vendor schedules. This article does not establish a universal historical trend or forecast.

Re-run the cost and quality model after material price, model, contract, or workload changes. Do not assume a currently unprofitable workflow will become profitable, or that future list-price reductions will rescue an inefficient design.

An illustrative twelve-week cost-optimization sequence

For a team starting from “we have an AI feature, costs are higher than expected,” the sequence below is a planning example. Change the duration and stop conditions to match the workload, evidence, and operational capacity.

Weeks 1-2: Measure.

  • Instrument per-call costs.
  • Build per-feature, per-user dashboards.
  • Identify the biggest cost contributors.

Weeks 3-4: Quick wins.

  • Trial prompt caching only where traces show repeated eligible prefixes and current provider rules fit.
  • Restructure the highest-cost eligible prompts, then measure cache hit rate, latency, quality, and billed cost.
  • Set each API’s current output cap only where measured task needs support it; test truncation and retries.
  • Implement budget alerts.

Weeks 5-6: Routing.

  • Identify simple tasks currently on flagship.
  • Build router for the 3-5 most-called endpoints.
  • Test for quality regression.

Weeks 7-8: Output and caching.

  • Constrain output lengths where not user-visible.
  • Add application-level response cache for common queries.
  • Add pre-filters for the highest-volume flows.

Weeks 9-10: Advanced.

  • Batch API for non-real-time work.
  • Provider alternatives evaluated.
  • Embedding cache, retrieval cache.

Weeks 11-12: Hardening.

  • Budget guards on every feature.
  • Cost dashboards in regular team review.
  • Documentation of patterns for future features.

At the end of the improvement cycle, publish the measured cost change and quality evidence. A schedule does not guarantee a savings percentage.

Measure first, then compound the savings

LLM costs can often be reduced, but the percentage and quality impact are workload-specific. Candidate techniques include caching, routing, output control, batching, pre-filtering, response caching, model selection, and budget guards.

Apply changes sequentially so their effects remain attributable; interactions may compound, overlap, or cancel out.

Use net contribution after inference, tooling, human review, infrastructure, maintenance, and support to decide whether the feature is economically sustainable.

Measure first. Optimize the biggest contributors. Maintain quality monitoring. Build cost discipline into the team’s regular work.

The result: AI features that scale economically, not just technically. That’s what makes AI a sustainable part of a product, not just a launch headline.

Read next

Continue through the same learning path with the next practical articles.