Using one model for every call can be unnecessarily expensive, but routing is not automatically an improvement.
A team may discover that classification, extraction, drafting, and hard reasoning have different quality and latency requirements. Routing can exploit those differences. It can also add a classifier call, provider failure modes, inconsistent safety behaviour, and more evaluation work. The only defensible savings claim is the one calculated from your measured tokens, current provider prices, and quality gates.
This is multi-model orchestration: selecting a model or specialised service for a defined call type under explicit quality, latency, privacy, and cost constraints.
This article covers the patterns, the routing logic, the trade-offs, and a step-by-step guide to implementing it.
Why one model isn’t optimal
Provider catalogues change frequently, but a workload can still define functional tiers:
High-capability reasoning tier: candidate models for difficult planning or analysis. Evaluate accuracy, tool behaviour, tail latency, and total generated reasoning tokens on your own cases.
General tier: candidates for user-facing drafting and mixed knowledge work where quality matters but extended reasoning may not.
Low-latency tier: candidates for bounded classification, extraction, and rewriting. Smaller does not guarantee adequate accuracy or lower end-to-end latency.
Local or small-model tier: candidates when data locality, offline operation, or marginal serving cost matters. Include hardware, operations, energy, concurrency, and quantisation effects in the comparison.
Specialised services (embedding, reranking, vision, speech): compare them with general models on the exact operation; specialisation is not proof of better quality or lower total cost.
Use live model and pricing pages at decision time: OpenAI models and pricing, Anthropic models and pricing, and Google models and pricing. Cache neither availability nor prices in a long-lived architecture decision.
A typical AI app makes many different types of LLM calls. Each call has its own requirements:
- Classifying user intent: evaluate a low-latency candidate against a labeled set, including ambiguous and out-of-scope cases.
- Extracting structured data: measure field-level accuracy and schema validity, not model size.
- Producing a user-facing response: measure factuality, policy compliance, and tail latency.
- Summarising past conversation: test omission of decisions, names, constraints, and negation.
- Background batch processing: measure throughput, retry cost, and deadline completion.
Do not assume the cheapest candidate is sufficient or that the most expensive candidate is best. Establish that with the same eval set for every route.
Calculate savings from telemetry
For each call type, collect:
- monthly call count;
- input, cached-input, and output token distributions;
- retry and cascade rate;
- tool, search, batch, or hosting charges;
- latency percentiles; and
- pass rate on the route’s release eval.
Calculate each candidate route with current prices:
monthly route cost = calls × (
input_tokens × input_price
+ cached_tokens × cached_price
+ output_tokens × output_price
) + tool_charges + hosting + expected_retry_cost
Compare that with the baseline only for candidates that pass the same quality and safety release criteria. Report the assumptions beside the result. A route that saves token cost but increases manual correction, incidents, or latency may cost more overall.
If telemetry is unavailable, run a shadow evaluation first. Do not publish a savings percentage from a generic traffic split.
The basic orchestration patterns
A few patterns recur in production multi-model systems:
Pattern 1: Task-based routing
Different types of tasks go to different models. This is the simplest pattern.
# Resolve these aliases from reviewed configuration; provider IDs change.
def route_request(task_type):
if task_type == "classification":
return SMALL_TIER
elif task_type == "extraction":
return EXTRACTION_TIER
elif task_type == "summarization":
return SUMMARY_TIER
elif task_type == "user-facing-response":
return RESPONSE_TIER
elif task_type == "complex-reasoning":
return REASONING_TIER
Tasks are classified by the calling code (it knows what it’s asking for). The routing is deterministic and easy to debug.
Pattern 2: Complexity-based routing
The system estimates the complexity of each request and routes accordingly.
def route_by_complexity(request):
complexity = estimate_complexity(request)
if complexity < 3:
return "small"
elif complexity < 7:
return "mid"
else:
return "flagship"
The complexity estimate can be heuristic (request length, keyword detection) or model-based (a cheap classifier scores the request). This pattern handles cases where the same task type varies in difficulty.
Pattern 3: Cascade routing
Try a cheap model first. If the output is good, use it. If not, escalate to a more expensive model.
def cascade(request):
cheap_response = call_model("small", request)
if is_acceptable(cheap_response):
return cheap_response
return call_model("flagship", request)
This works when “acceptable” is detectable — by confidence scores, validators, or a separate quality-check LLM. It’s powerful: most simple requests get answered by the cheap model; only the hard ones reach the expensive one.
Pattern 4: Specialty routing
Use specialised models for specialised tasks:
- Embeddings: use a dedicated embedding model (much cheaper than a chat model used to embed).
- Reranking: use a dedicated reranker.
- Vision: use a vision-specialised model for image analysis.
- Voice: use a voice model for transcription/synthesis.
- Code: use a code-specialised model for code tasks.
Specialised or smaller models may be faster, cheaper, or better on a narrow task, but none of those advantages follows from the label alone. Benchmark the exact model, provider, prompt, language, latency percentile, price sheet, and evaluation set.
Pattern 5: Provider routing
Use models from multiple providers for redundancy and pricing leverage.
providers = ["openai", "anthropic", "google"]
preferred = "anthropic" # primary
fallback = "openai" # fallback
def call_with_failover(request):
try:
return call(preferred, request)
except (RateLimit, ProviderError):
return call(fallback, request)
This gives resilience against single-provider outages and rate limits. It also lets you take advantage of pricing changes — when a provider cuts prices, shift more traffic there.
A realistic example: customer support AI
To make it concrete, here’s how a customer support AI might use multi-model orchestration.
The system has these steps per ticket:
Step 1: Classify intent. What is the customer asking about?
Routing candidate: a low-latency model that passes the labeled intent set.
Step 2: Determine urgency and sentiment. Is the customer frustrated? Is this urgent?
Routing candidate: the same model only if urgency false negatives meet the separately defined safety threshold; sentiment is not a reliable proxy for urgency.
Step 3: Retrieve relevant knowledge.
Routing candidate: embedding retrieval plus a reranker, evaluated on a ticket-specific retrieval set.
Step 4: Determine if the AI can answer this or needs human escalation.
Routing candidate: a model evaluated specifically for escalation recall. Rules should force human escalation for account access, safety, legal, financial, or other policy-defined cases.
Step 5 (if AI can answer): Generate the customer-facing response.
Routing candidate: a higher-quality general model. Keep it draft-only until factuality, policy, privacy, and tone pass release criteria.
Step 6 (if AI cannot answer): Generate a summary for the human agent.
Routing candidate: a lower-cost model whose summaries preserve the issue, evidence, attempted steps, customer constraints, and uncertainty.
Step 7: Quality check. Did the response meet our standards?
Routing candidate: deterministic checks plus a calibrated judge. A judge model is not an independent guarantee; sample its decisions with humans.
Instrument each step, then fill in the cost formula above with measured token distributions, current prices, escalation share, retry share, and human-review cost. This example intentionally provides no generic savings estimate: ticket mix and acceptance thresholds determine the result.
The routing logic
A few approaches to implementing routing:
Approach 1: Hard-coded by task type
Simplest. You know what task you’re calling, you pick the model.
def classify(text):
return openai_client.chat.completions.create(
model=SMALL_TIER, # your provider's current small model
messages=[{"role": "user", "content": f"Classify: {text}"}],
)
def respond(context, query):
return claude_client.messages.create(
model=RESPONSE_TIER, # reviewed config alias, not a frozen provider ID
max_tokens=1024,
messages=[{"role": "user", "content": f"Context: {context}\n\nQuery: {query}"}]
)
Pros: Transparent, easy to debug, easy to change. Cons: Doesn’t adapt to request complexity within a task type.
Approach 2: Router model
A small model classifies each request and routes it.
ROUTER_PROMPT = """
Classify this request as: trivial, moderate, or complex.
Output one word.
Request: {request}
"""
def route(request):
classification = small_model_call(ROUTER_PROMPT.format(request=request))
return MODEL_BY_COMPLEXITY[classification]
Pros: Adapts to complexity within a category. Cons: Adds latency (the router call), adds a failure point, requires tuning.
Approach 3: Embedding-based router
For requests that fall into known patterns, use embedding similarity to past examples.
def route(request):
embedding = embed(request)
closest = find_nearest_example(embedding)
return closest.suggested_model
Pros: Fast (just a vector lookup), gets smarter with more data. Cons: Requires building a labeled set of examples.
Approach 4: Cascade
Try cheap first; escalate if needed.
def cascade(request):
cheap = small_model_call(request)
if validates(cheap):
return cheap
return flagship_call(request)
Pros: Adaptive, low average cost. Cons: Slow for cases that need escalation (two calls), requires reliable validation.
In practice, many production systems use a hybrid: hard-coded routing for the main task types, with cascades for specific high-variance subtypes.

The pitfalls
A few mistakes to avoid:
Pitfall 1: Optimising for cost while degrading quality
It’s easy to route everything to small models and watch the cost drop. It’s harder to notice that quality also dropped. Always pair routing changes with quality monitoring.
A useful discipline: shadow or A/B test a cheaper route until the sample covers important input classes and failure modes. A fixed week is not evidence by itself. Do not ship the change unless the predeclared quality and safety criteria hold.
Pitfall 2: Over-engineering the router
A router with many task types and opaque logic can be harder to maintain than the routing it replaces. Start with the smallest number of routes your measurements justify.
Add a route only when it changes an operational decision and improves a measured constraint enough to justify its ownership, tests, and fallback path.
Pitfall 3: Ignoring latency
Lower-priced models are not necessarily faster. Measure latency and quality separately. A cascade (try one model, then fall back) adds at least one extra attempt for fallback cases and can materially increase tail latency in user-facing flows.
For a user-facing latency-sensitive response, evaluate a direct route against a cascade on both pass rate and tail latency. Cascades are often easier to tolerate in batch or asynchronous work, but the correct route is workload-specific.
Pitfall 4: Not handling provider failures
When you depend on multiple models, you have multiple ways to fail. A flagship model goes down, a rate limit kicks in, an API key expires. Your routing logic needs fallbacks.
Define explicit behaviour for rate limits, timeouts, and provider errors. A cross-provider fallback is appropriate only if its data-processing terms, regional path, tool/schema contract, and eval results are acceptable. Otherwise fail closed, queue, or escalate; returning a materially worse answer is not availability.
Pitfall 5: Not measuring per-route quality
You need to know which route is doing well and which isn’t. This means evaluation, ideally automated.
A useful setup: for every production call, log the model used, the request, the response, and (where possible) some quality signal (user feedback, downstream metrics, automated eval). Roll up per-route metrics. Catch quality drift before users complain.
Revalidate live dependencies
Model IDs, retirement dates, context limits, prices, rate limits, regional processing, and structured-output/tool semantics can change independently. Review live provider documentation and rerun route evals before changing an alias. An OpenAI-compatible transport does not guarantee equivalent schemas, tool behaviour, token accounting, safety policy, or data handling.
A starter checklist
If you’re building a multi-model system from scratch, or migrating from single-model:
-
Map your tasks. What types of LLM calls does your app make? Roughly how often? Roughly how expensive?
-
Categorise by complexity. For each task type, decide: trivial, moderate, or complex. Match to model tier.
-
Build a router. Start with hard-coded task-based routing. Don’t over-engineer.
-
Define failure behaviour. Use a tested fallback, queue, fail-closed response, or human escalation according to consequence and data policy.
-
Measure quality per route. Set up logging and basic evaluation. You need to know if quality holds.
-
Iterate. Move tasks to cheaper models where quality holds. Move tasks back to expensive models where quality breaks. Adjust over time.
-
Don’t stop tuning. Models change. New ones launch. Prices shift. A routing setup that’s optimal in May 2026 may be suboptimal in November 2026.
Route only where evidence supports it
Multi-model orchestration can reduce cost or latency when call types genuinely differ and each route is evaluated. It can also increase operational cost and lower consistency. Publish the measured baseline, routed result, quality gates, and sample period instead of a universal savings claim.
The routing conditional may be short; the production work is configuration control, evaluation, observability, privacy review, retry semantics, and fallbacks.
Map your tasks, benchmark viable candidates, route only where the evidence justifies it, and keep measuring after release.



