If you’re building real AI products in 2026, you no longer ask “should I use OpenAI or Anthropic?” That framing is two years stale. You’re making dozens of decisions across a layered stack, and most of them matter.
Here is a working architecture map: what belongs at each layer, which trade-offs to measure, and which volatile choices must stay replaceable. Product names below were checked against vendor documentation on 4 August 2026. They are examples, not a ranking or a recommendation for a workload we have not evaluated.
Primary references for the volatile model layer: OpenAI’s current model catalogue, Anthropic’s model overview, Google’s Gemini model catalogue, and DeepSeek’s model and pricing page.
The layers
The stack, roughly:
┌────────────────────────────────────┐
│ Application Layer │ Your product / agent / workflow
├────────────────────────────────────┤
│ Orchestration / Frameworks │ LangGraph, CrewAI, custom, direct
├────────────────────────────────────┤
│ Prompt + Context Management │ Prompt templates, context engineering
├────────────────────────────────────┤
│ Retrieval / Memory │ RAG, vector stores, structured memory
├────────────────────────────────────┤
│ Tool / MCP Layer │ Tool calling, MCP servers, function APIs
├────────────────────────────────────┤
│ Model Layer │ Specific model selection, routing
├────────────────────────────────────┤
│ Inference Layer │ Hosted APIs, self-hosted, edge
├────────────────────────────────────┤
│ Observability / Evals │ Logging, tracing, eval suites
└────────────────────────────────────┘
Each layer has multiple viable options. The choice at one layer constrains options at others. Decisions early on are sticky — model choice influences inference choice influences orchestration choice.
We’ll go through each.
Layer 1: The model layer
Models in 2026 cluster into rough tiers, and selecting from each is the most consequential per-call decision in your system.
Capability-first tiers. Current examples include OpenAI GPT-5.6 Sol, Claude Fable 5 / Opus 5, Gemini 3.1 Pro Preview, and DeepSeek-V4-Pro. Their reasoning controls, tool support, latency, and prices differ. Do not infer that the model with the highest list price wins your task; compare pinned versions on a representative eval set.
Balanced tiers. Current examples include GPT-5.6 Terra, Claude Sonnet 5, and Gemini 3.5 Flash. These are candidates for general production traffic, not universal defaults. Measure task success, time to first token, total latency, and cost using your actual input and output distribution.
Efficiency tiers. Current examples include GPT-5.6 Luna, Claude Haiku 4.5, Gemini 3.5 Flash-Lite, and DeepSeek-V4-Flash. They are useful candidates for classification, extraction, routing, and high-volume generation when an eval shows the lower-cost model still clears the acceptance bar.
Small and on-device models. Small open-weight models can be effective for narrow, structured tasks and privacy-sensitive local processing. Their speed depends on hardware, quantization, sequence length, and concurrency; benchmark instead of publishing a universal latency claim.
(Availability and names verified 2026-08-04. Prices are intentionally not copied into this overview: use the OpenAI, Anthropic, Google, and DeepSeek pages in a dated cost model.)
Specialised models. Embedding, reranking, vision, voice, and code-specialised models may provide a better cost, latency, or quality point for a narrow operation. Compare them with the general-model baseline on that operation; neither specialization nor a lower list price proves a better system result.
Open-weight frontier. DeepSeek V4, Qwen 3, the Llama family, and Mistral models can be hosted by a provider or on infrastructure you operate, subject to each model’s licence. “Open source” is not a safe blanket label: weights, training code, dataset transparency, and commercial terms vary.
The implications:
- One model does not fit every call. Routing is worthwhile only when the measured savings exceed router errors and operational complexity.
- The frontier is moving every quarter. Build to swap models, not lock in.
- Open-source is now genuinely viable for many production use cases, not just experiments.
Layer 2: The inference layer
Where does your model actually run?
Closed API providers — OpenAI, Anthropic, Google. Fastest path to running, best models, most reliable. You pay a premium and accept the data/security model.
Open-source inference providers — Groq, Together AI, Fireworks, Replicate. Run open models with various tuning. Often much faster than self-hosting; competitive prices. (This market consolidates fast — several 2024-era providers no longer exist; verify any vendor before committing.)
Cloud-native — AWS Bedrock, Azure OpenAI, Google Vertex. Wraps closed and open models with your cloud’s auth, billing, compliance. Necessary for many enterprise contexts.
Self-hosted — vLLM, TGI, SGLang, LMDeploy on your own GPUs. Potentially lower marginal cost at sustained utilization, with substantially higher operational complexity. Calculate break-even from your traffic shape, hardware or rental cost, utilization, staffing, and availability requirements; there is no universal monthly-spend threshold (see self-hosted vs hosted inference).
Edge / on-device — Apple Intelligence, MediaPipe, ONNX, GGUF models via Ollama or llama.cpp. Free per call, but constrained model capability. Increasingly viable for narrow use cases.
The trade-offs:
- Latency matters: voice agents, conversational UX need fast first-token. Groq, Cerebras, and on-device dominate here.
- Throughput matters for batch: if you process millions of records, you want high-throughput, not low-latency.
- Legal and contractual requirements matter: GDPR roles and transfer rules, regulated-data obligations, customer contracts, and data-residency commitments can constrain providers and regions. A SOC 2 report is assurance evidence, not a law or an automatic permission to process data.
- Vendor risk matters: depending solely on one provider is a single point of failure. Multi-provider is good hygiene.
A common 2026 pattern: hosted closed models for the highest-quality user-facing requests, hosted open-source for high-volume cheaper work, on-device for narrow latency-sensitive features. Self-hosting only when scale and economics justify the operational burden.
Layer 3: Tools and MCP
LLMs alone can’t do much. They become useful when they can call tools — functions you define that give them access to data, APIs, and actions.
Native function calling. Every major model supports a structured function-calling API. You define functions with JSON schemas; the model decides when to call them; you execute the call; you return results.
MCP (Model Context Protocol). A standardized protocol for clients to connect to tool and context servers. It can reduce vendor coupling at the protocol boundary, but authentication, authorization, deployment, and client-specific behavior still need testing. Start with the current MCP specification, not a copied tutorial.
Direct integrations. For high-volume specific use cases (e.g., specific CRM, specific database), often easier to write a direct adapter than a generic MCP server.
MCP has broad cross-vendor adoption, but “use MCP for every integration” is not an engineering rule. Choose it when multiple clients need the same capability or protocol portability matters. A direct typed adapter can be simpler for one high-volume internal path.
A few implementation realities:
- Tool descriptions matter enormously. A poorly described tool will not be used correctly. Tool docstrings should be written like prompts.
- Tool count matters. Models with 50+ tools available perform worse than ones with 5-10 relevant tools. Curate aggressively.
- Error handling matters. Tool errors need to be communicated to the model in a structured way so it can adapt.
- Authorization is hard. A multi-user system where the LLM has different permissions for different users is non-trivial. Don’t let the LLM make authorization decisions; do it in the tool wrapper.
Layer 4: Retrieval and memory
LLMs need data they weren’t trained on. This is the retrieval layer.
Vector databases. Pinecone, Weaviate, Qdrant, Chroma, PostgreSQL with pgvector, Turbopuffer. Stores embeddings; serves nearest-neighbor queries. Mature, well-understood. The default for semantic retrieval.
Hybrid search. Combines vector search with traditional BM25 keyword search. Catches both semantic and lexical matches. Use Reciprocal Rank Fusion to combine. Tools: Elasticsearch, OpenSearch, Vespa.
Knowledge graphs. Neo4j, Memgraph, custom triple stores. For data with rich relationships. Used in graph RAG architectures. More work to build, often higher quality for relationship-heavy domains.
Specialised RAG platforms. LlamaIndex (now mature), LangChain RAG abstractions, Haystack. Higher-level frameworks for common patterns.
Reranking. Cohere Rerank, Voyage, custom cross-encoders. After initial retrieval, rerank the top candidates with a more expensive model. The effect depends on the first-stage retriever, candidate depth, corpus, and metric; keep it only when an evaluation set shows a useful quality gain.
Memory. For agents and conversations, structured memory layers — Mem0, Letta (formerly MemGPT), or custom. Distinguish short-term (current conversation), medium-term (recent topics), long-term (durable facts about the user/account).
The architectural question: where does this layer live?
- In-app: the LLM call is wrapped in retrieval logic written by your team.
- At the MCP layer: retrieval exposed as tools.
- As a service: a dedicated retrieval service your apps call.
For monolithic single-product systems, in-app is fine. For multi-product organizations, treating retrieval as a service (with consistent quality and policy) pays off.
Layer 5: Prompt and context engineering
In 2026, “prompt engineering” is mostly synonymous with “context engineering” — managing what goes into the context window for each call.
The components:
Prompts. Often templated with variables. Stored in version control. Tested with eval suites. Treated like code.
Prompt management. Tools like Promptfoo, Langfuse, PromptLayer, or in-house systems. Versioning, A/B testing, rollback. (Helicone and similar LLM proxies belong in the observability layer below, not here — easy to conflate the two categories.)
Context strategy. Decisions about what to include in each call:
- System prompt (stable, defines behavior).
- Retrieved knowledge (dynamic, from RAG).
- Conversation history (managed, often summarized at length).
- Few-shot examples (chosen dynamically based on the query).
- Tool descriptions (filtered to relevant tools only).
- The user’s current query.
Context compression. As context grows long, the model degrades. Strategies: summarize old turns, extract key facts to a structured memory, prune irrelevant content. Active research area.
Long-context use. Several current flagship families advertise roughly one-million-token windows, including GPT-5.6, current Claude flagships, and Gemini models. Capacity is not retrieval quality. Long-context research such as Lost in the Middle shows that relevant information can be used unevenly by position, and results vary by model and task. Evaluate the exact model, prompt, document order, and context length you plan to ship.
Layer 6: Orchestration
How do you coordinate multi-step LLM workflows and agents?
Direct API. Just write the loop yourself in Python or TypeScript. Best for simple cases and for understanding what’s actually happening.
LangChain / LangGraph. Widely used. LangGraph (state machine for agents) has matured significantly. Heavy abstractions, learning curve, but powerful.
CrewAI. Multi-agent framework focused on role-based agents. Easier to start than LangGraph; less flexible.
LlamaIndex agents. A candidate when the team already uses its data and retrieval abstractions; verify current APIs and compare it with a smaller direct implementation.
OpenAI Agents SDK. An OpenAI-oriented orchestration option; compare it with a direct Responses API loop before accepting the framework dependency.
Claude Agent SDK. Anthropic’s agent runtime for Claude-oriented coding and tool workflows; it is not a generic synonym for the Claude API.
Custom. For mature teams shipping production agents, custom orchestration is common — frameworks impose costs (abstraction tax, debugging complexity, version churn) that outweigh benefits.
There is no evidence-based rule that every framework prototype should be rewritten. Start with the smallest abstraction that expresses your workflow, instrument it, and migrate only when framework constraints or operational cost are measured.
Layer 7: Observability
You cannot ship serious LLM applications without observability. Every production system needs:
Tracing. Every LLM call captured: timestamp, model, input, output, latency, cost, success/failure. Trees for multi-step traces.
Cost tracking. Per-call, per-feature, per-user. Costs are large and unbounded; without tracking you find out at the end of the month.
Quality monitoring. Automated quality checks on a sample of production traffic. Alerts on quality drops.
User feedback capture. Thumbs up/down, explicit feedback, implicit signals (retry rate, abandonment).
Debugging. When something breaks, you need to see the full call chain. A failed agent run has many possible failure points.
Tools: LangSmith, Helicone, Arize, Phoenix, Braintrust, Weights & Biases, Datadog LLM Observability. Each has different strengths; pick one early and stick with it.
For small teams: even a simple Postgres table with one row per LLM call gets you 80% of what you need. Move to a tool when scale or feature needs justify it.

Layer 8: Evals
The single most important layer for serious production work.
Offline evals. A defined dataset; expected outputs; scoring. Run before deploying changes. Catches regressions. (Practitioner framing: evals for non-engineers; deeper patterns in evals that catch regressions.)
Online evals. Sample of production traffic scored automatically (LLM-as-judge) or via user signals. Catches drift.
Pre-deployment evals. Before any prompt or model change goes live, eval suite runs and is reviewed. Becomes part of CI.
Eval taxonomy. Different evals for different concerns:
- Behavioral: does it do what we expect?
- Safety: does it refuse what we want refused?
- Quality: how good is the output?
- Robustness: how does it handle adversarial inputs?
- Cost/latency: are we within budget?
Tools: Promptfoo, Braintrust, LangSmith, custom suites. All have a place; Promptfoo is the easiest start.
Layer 9: The application layer
This is where your specific product lives. The decisions here:
Agent vs workflow. Agents (LLM in a loop with tools) are powerful but harder to make reliable. Workflows (fixed sequence of LLM calls) are easier and often sufficient. Default to workflows; reach for agents when truly needed.
Synchronous vs asynchronous. User-facing real-time? Batch background? Streamed? Affects model choice, infrastructure choice, UX design.
Single-tenant vs multi-tenant. Customer-specific data isolation requirements drive significant architecture decisions.
On-prem vs cloud. Compliance, security, or cost may push you on-prem. Operational complexity is much higher.
Edge cases. Hallucinations, prompt injections, abuse. Production systems need guardrails. Don’t ship without them.
Trade-offs that matter
A few trade-offs worth being explicit about:
Quality vs cost vs latency
These objectives often conflict, but they are not a fixed “pick two” law. Plot task success, tail latency, and complete cost for the evaluated candidates and choose against explicit constraints.
- High quality + low latency = expensive.
- Low cost + low latency = lower quality.
- High quality + low cost = high latency (batch processing, or reasoning models).
Pick your priorities per task. Don’t optimize all three; that path leads to mediocrity in all.
Build vs buy
For each layer, you can build or buy.
- Build: more control, more maintenance, more cost (engineer time), differentiating capabilities.
- Buy: faster start, less control, ongoing vendor risk, undifferentiating capabilities offloaded.
A good heuristic: buy the commodity layers (vector storage, basic observability), build the differentiating layers (your specific orchestration, your prompts, your evals). Reversing this — buying your differentiation and building your commodity infrastructure — is a common mistake.
Open-source vs closed
Open-weight models are viable candidates for many tasks, especially when deployment control matters. Whether one is faster, cheaper, or sufficiently capable depends on the exact model, serving path, hardware, concurrency, and task evaluation; do not infer it from the licence category.
The decision factors:
- Quality requirements. Compare pinned candidates on the task and critical slices; licence or access model does not determine quality.
- Cost at scale. Compare current API bills with benchmarked, availability-adjusted hosting TCO at the forecast load shape.
- Privacy/compliance. Choose a deployment boundary from data classification, contracts, architecture, and qualified review; self-hosting is not automatic compliance.
- Customization. Verify whether the candidate’s licence and provider support the needed adapters, training, constrained output, tools, or serving changes.
- Operational capacity. Managed APIs and self-hosting have different security, availability, migration, incident, and vendor-dependency burdens; neither is trivial for every workload.
Hybrid routing is one candidate when multiple model/deployment paths create a measured benefit that exceeds router errors, policy complexity, and operating cost. This article has no representative deployment survey proving it is the majority pattern.
Latency vs reasoning depth
Reasoning controls and model tiers can change task quality, latency, and billed usage. Compare the current pinned APIs and measure the full latency distribution; a “reasoning” label does not prove an improvement on your hard cases.
A pattern: route simple queries to fast models, hard queries to reasoning models. Use a router (small model or heuristic) to decide.
Long context vs RAG
You can stuff context into the model (using a million-token window) or you can retrieve relevant chunks (using RAG).
- Long context: avoids a retrieval index but still needs document ordering, access control, token/cost limits, and position/distractor evaluation.
- RAG: adds ingestion and retrieval infrastructure and may reduce supplied context; its quality and total cost are empirical.
The defensible answer comes from comparison: long-context baseline, retrieval baseline, and where relevant a hybrid. Evaluate answer correctness, evidence recall, citation quality, latency, cost, permission enforcement, and freshness on the same workload.
Agents vs workflows
Discussed above. Default to workflows; use agents when you genuinely need the flexibility. Many “agent” systems we see should be workflows.
A 2026 reference architecture
To make all this concrete, here’s what a typical production system looks like for a mid-sized SaaS product with AI features:
User → Application (React/Next.js)
↓
API gateway / auth
↓
LLM Service (your wrapper)
↓
Router (small model or heuristic)
├→ Simple tasks: an evaluated efficiency-tier model
├→ Standard tasks: an evaluated balanced-tier model
├→ Hard tasks: a pinned capability-tier model
└→ Special: vision/voice/embedding specialists
↓
Tool layer (MCP servers + direct integrations)
↓
Retrieval layer (Pinecone + hybrid + reranker)
↓
Observability (Helicone or LangSmith)
↓
Eval suite (Promptfoo, runs in CI)
For this reference architecture, calculate cost per completed workflow from traced tokens, tool and retrieval calls, cache state, retries, infrastructure, and human review. Estimate engineering and operating effort from a scoped backlog and the team’s observed delivery rate. No transferable per-user price or delivery duration is asserted here.
What commonly goes wrong
Failure patterns that recur across production LLM stacks:
Pattern 1: Single model for everything. Cost overruns, quality issues. Fix: routing.
Pattern 2: No observability. Can’t debug, can’t measure, can’t improve. Fix: instrument early.
Pattern 3: No evals. Quality drifts unnoticed. Fix: evals from day one.
Pattern 4: Framework lock-in. LangChain or CrewAI debugging becomes a full-time job. Fix: don’t use frameworks unless they save more than they cost. Rewrite to direct code when patterns are clear.
Pattern 5: Building infrastructure that should be bought. Custom vector DB? Probably wasted time. Custom observability? Probably wasted time. Buy the commodity layers.
Pattern 6: Buying infrastructure that should be built. Outsourcing your prompts to a third party. Outsourcing your evals. These are your competitive moat; own them.
Pattern 7: Ignoring prompt injection. Production system without input sanitization for user-provided content. Big risk; mitigate early.
Pattern 8: Trusting agents in high-stakes flows. A LangGraph agent that authorizes refunds with no human review. This will eventually go wrong. Add human-in-loop for consequential actions.
Pattern 9: Optimizing for the wrong thing. Optimizing inference cost when total cost is dominated by engineering time. Or optimizing latency when users don’t notice. Measure what actually matters.
Pattern 10: No multi-provider plan. When (not if) your primary provider has an outage, you’re down. Have a fallback configured.
Three bets, dated (check back mid-2027)
Predictions are cheap; dated, falsifiable ones are not. Three we are willing to be wrong about in public:
-
EU-hosted inference reaches practical price parity for a defined mid-tier workload by mid-2027. “Parity” means no more than a 10% premium for the same pinned model, throughput, availability target, and support tier. If that happens, EU-region processing becomes the default candidate for relevant workloads; legal suitability still depends on the complete processing arrangement. Confidence: moderate.
-
The framework layer keeps consolidating while protocol boundaries remain more durable. We will count this as supported if at least two major agent frameworks merge or retire by mid-2027 while MCP remains implemented by multiple independent client vendors. We therefore keep tool contracts portable and orchestration replaceable. Confidence: high.
-
Small-model routing stops being an optimization and becomes the default architecture. If flagship prices hold while small tiers keep improving, “flagship for everything” will read the way “bare metal for everything” reads to a cloud engineer. Confidence: high for cost-sensitive SMEs.
What we deliberately do not predict: model rankings. Any specific ranking printed here would be stale before this page’s next review date.
Get the architecture right
The 2026 LLM stack is real, layered, and the choices matter. The teams winning are those who:
- Understand the full stack, not just the parts they touch.
- Make explicit trade-offs (quality, cost, latency) per call.
- Build the parts that differentiate; buy the parts that don’t.
- Instrument from day one (observability, evals).
- Stay nimble (model-portable, multi-provider).
The teams losing are those who picked one vendor, hard-coded its API, never instrumented, never measured, and now find themselves with a system that’s expensive, brittle, and impossible to improve.
Get the architecture right. Everything else gets easier.



