Observability for LLM apps: tracing, costs, latency, quality drift
Advanced12 min readAI for Business

Observability for LLM apps: tracing, costs, latency, quality drift

Extend ordinary observability with multi-step traces, attributable cost, prompt and model versions, evaluated quality signals, privacy controls, and workload-derived alerts.

What you should be able to do

Keep ordinary service telemetry, then add model, prompt, retrieval, tool, cost, and evaluated-quality context. Capture raw prompts or outputs only for a justified purpose with access, minimisation, retention, and deletion controls.

Saved only in this browser.
In this article

LLM applications retain ordinary software failure modes and add probabilistic output, model and prompt drift, tool behavior, retrieval quality, and usage-sensitive cost. Some failures produce stack traces; others appear only in evaluated outputs, user reports, or changed distributions.

Traditional observability (Datadog, New Relic, Sentry) tells you the API call succeeded in 8.4 seconds and used 12,847 input tokens. It doesn’t tell you whether the response was good, whether the model hallucinated, whether the wrong tool was called, or whether quality has drifted since last week.

Production LLM systems need additional semantic data on top of ordinary traces, metrics, and logs. Start with the OpenTelemetry generative-AI semantic conventions and extend only where your product needs more detail. This article defines an illustrative event shape; AIExpert has not published a production trace or dashboard proving every field below.

What’s different about LLM observability

A few characteristic features of LLM systems that traditional observability doesn’t address:

Non-determinism at the unit level. The same input produces different outputs across calls. “Did this work right?” can’t be answered by checking status codes.

Quality as a primary metric. Latency and cost matter, but quality matters most — and is the hardest to measure.

Multi-step traces. A user query can trigger multiple model, retrieval, and tool calls. Each operation belongs to a larger trace.

Usage-dependent cost. Cost varies with model, token counts, caching, tools, and provider-specific charges. Attribute actual billed usage at call or batch level rather than assuming a universal per-call range.

Drift over time. Models update. Prompts evolve. Input distributions shift. Quality moves; you need to see that movement.

Sensitive payloads. The inputs and outputs are often the most valuable diagnostic data but also the most sensitive. Logging discipline matters.

Long async flows. Agent runs that take minutes. Background batch jobs. Streaming responses. Traditional request/response observability doesn’t fit.

These are failure modes the instrumentation should be designed to reveal; which ones dominate must be established from your own traffic.

The observability stack

A practical LLM observability design can include these layers, selected by workload and risk:

1. Call-level instrumentation. Every LLM call emits approved operational metadata such as latency, billed usage, model/revision, and status. Raw input or output is captured only when separately justified and controlled.

2. Trace-level instrumentation. Multi-call workflows are stitched into traces. You can see the full call chain for a user request.

3. Application-level metrics. Per-feature, per-user, per-tenant aggregations.

4. Quality monitoring. Sampled or full automated quality assessment.

5. User feedback capture. Explicit (thumbs up/down) and implicit (regenerate, abandon) signals.

6. Alerting. Real-time alerts on cost spikes, latency degradation, quality drops, error rate increases.

7. Debugging tools. When something breaks, authorized responders can inspect the minimum approved evidence needed to understand it; raw inputs and outputs are not guaranteed or universally available.

We’ll go through each.

Call-level instrumentation

Every LLM call should produce an approved metadata record. The commented payload and identity fields below are optional sensitive fields, not defaults:

{
  "call_id": "uuid",
  "timestamp": "2026-08-04T14:23:45Z",
  "trace_id": "uuid",  // for grouping into traces
  "span_id": "uuid",   // for parent-child relations
  "feature": "summarize_document",
  "prompt_version": "v3.2",
  "model": "provider-model-revision",
  "provider": "anthropic",
  "input_messages": null,  // optional; approved sampled/redacted payload only
  "output_message": null,  // optional; approved sampled/redacted payload only
  "input_tokens": 1842,
  "output_tokens": 384,
  "total_tokens": 2226,
  "cost_usd": 0.0084,
  "latency_ms": 2340,
  "first_token_ms": 1240,  // streaming
  "status": "success",
  "error": null,
  "subject_ref": "pseudonymous_ref",  // optional, purpose-specific lookup
  "tenant_ref": "tenant_scoped_ref",
  "metadata": {
    "session_id": "...",
    "experiment_arm": "v3_test"
  }
}

This is an illustrative application event, not a mandatory schema. Capture the minimum needed for a defined operational purpose. Raw prompts and outputs are opt-in sensitive payloads, not default telemetry.

Key implementation choices:

Where to log. Options:

  • A dedicated observability tool (Helicone, LangSmith, Phoenix, Braintrust, Arize).
  • A general observability platform with LLM extensions (Datadog LLM Observability, Sentry).
  • Your own logs/database.

Choose from requirements: OpenTelemetry export, trace and tool-call visualization, data location, self-hosting, redaction, access control, retention/deletion APIs, model-price maintenance, evaluation support, and total cost. A dedicated tool, an existing APM, or a small owned implementation can each be correct; test export and deletion before committing.

How to instrument. Options:

  • A proxy that sits between your app and the LLM provider (Helicone’s model).
  • A wrapping SDK in your application code.
  • A wrapper class that you call manually for each LLM invocation.

Proxies centralize instrumentation but add a network hop and failure domain. SDK wrapping keeps instrumentation in process but couples code to the integration. Manual wrappers can be precise but require coverage tests. Measure overhead and failure behavior for the selected approach.

A pragmatic approach: SDK wrapping at the boundary where your app calls the LLM. One place to instrument; everything else flows through.

What to log. Apply data classification and retention before enabling payload capture. Under GDPR, Article 5 data minimisation and storage-limitation principles still apply to observability data.

  • Prefer prompt version, hashes, token counts, policy decisions, and derived metrics over raw content.
  • If payload capture is justified, sample it, redact it before export, encrypt it, restrict access, and set a short retention period.
  • Do not keep an automatically retrievable raw original merely because a redacted copy was logged; that recreates the sensitive-data store and needs its own lawful purpose and controls.
  • Don’t log credentials, even in error paths.
  • For streaming, log both first-token-latency and total-latency.

Trace-level instrumentation

A single user action often involves many LLM calls. Without trace-level instrumentation, you have a thousand call logs and no way to know which calls were part of which user action.

Implementation:

Trace ID generation. Generate a unique trace ID at the start of a user request. Pass it through all subsequent calls.

Parent-child spans. Within a trace, each call has a span ID and (optionally) a parent span ID. This creates a tree showing the call hierarchy.

Operation naming. Each span is named (“summarize_document”, “extract_entities”, “tool_call:search”). The trace shows the full operation chain.

A trace view in your UI:

Trace abc-123 (12.3s total)
├─ classify_intent (450ms) [small-model-revision]
├─ retrieve_documents (1.2s) [embedding + search]
├─ generate_response (8.5s) [generation-model-revision]
│   ├─ tool_call: search_internal (320ms)
│   ├─ tool_call: lookup_customer (180ms)
│   └─ generate_final_text (7.5s)
└─ judge_response_quality (2.1s) [claude-haiku-4-5]

Now you can see what your system actually did for this user request. You can find slow calls, expensive calls, failed calls — in context.

Candidate implementations include dedicated LLM-observability products and general APM systems using OpenTelemetry. Verify trace propagation, tool-call representation, export, redaction, deletion, access control, and failure behavior in a representative trial rather than relying on a product list.

Application-level metrics

Beyond individual calls, aggregate metrics:

Per-feature.

  • Call volume.
  • Average latency.
  • p50, p95, p99 latency.
  • Average cost per request.
  • Error rate.
  • Quality score (if measured).

Per-user/per-tenant.

  • Calls per user per day.
  • Cost per user.
  • Heavy users / abuse patterns.

Per-model.

  • Volume by model.
  • Cost share by model.
  • Error rate by model.
  • Quality (where measured) by model.

Per-feature × per-model.

  • Which features use which models?
  • Where could we route to cheaper models?

These dashboards drive operational decisions: which features are expensive, which are slow, which need optimization.

Quality monitoring

The hardest layer: automated quality assessment.

For evals you run on a defined dataset. For online quality monitoring, you assess production traffic.

Approaches:

LLM-as-judge on a sample. Define a sample from traffic volume, risk slices, privacy constraints, detection target, and budget. Run a calibrated judge on eligible examples, preserve human adjudication for disputed or high-impact cases, and track the result by prompt/model version. Judge models have position, verbosity, and self-preference biases; they are measurements to validate, not ground truth.

Implicit signals. Track regenerations, abandonments, error rates, time-to-completion, follow-up message frequency. These are weak signals but cheap. Use as leading indicators.

Explicit user feedback. Thumbs up/down, “did this help?” buttons, explicit reports. Highest signal but lowest volume.

Pattern detection. Specific bad patterns (“I cannot help with that”, “I’m just an AI”, repeated refusals) flagged automatically. Catches some regressions immediately.

An implementation record should state the sampling rule, excluded data, rubric, judge version, human calibration set, uncertainty, aggregation window, alert threshold, and cost ceiling. Derive change thresholds from repeated baseline runs and the severity of missed regressions; a copied percentage has no statistical or business meaning for a different workload.

Cost observability

Retries, loops, unexpectedly long inputs or outputs, routing changes, and provider-price changes can raise cost quickly. Detect each mechanism directly instead of assuming a multiplier.

Cost observability layers:

Per-call cost. Every call’s cost computed at log time. Aggregations available immediately.

Budget alerts. Daily, weekly, or monthly budgets per feature and tenant, with warning and enforcement thresholds chosen from forecast variance and business criticality.

Anomaly detection. Compare cost and usage with the correct traffic-mix baseline, and separately cap pathological single-run loops, tokens, tools, retries, duration, and spend.

Cost attribution. Per-feature, per-tenant, per-user costs. Find the heavy hitters.

Forecast. Based on current trajectory, what will the bill be at end-of-month?

A useful dashboard view: a single panel showing today’s cost vs the rest of this week vs last week, broken down by feature.

Set budget limits from measured traffic and business criticality. Prefer staged controls—alert, throttle, downgrade a non-critical path, then circuit-break—because a universal “shut off at 10×” rule may react far too late or take down an essential workflow.

Latency observability

LLM latency is more nuanced than typical APIs:

Total latency. From request to final response.

Time-to-first-token (TTFT). For streaming, when does the user see the first character? This dominates perceived latency for chat UX.

Time-to-last-token (TTLT). How long until the response is complete?

Tokens-per-second. Output rate. Some models stream slower than others.

Tool call latency. For agent flows, time spent in tool calls vs LLM calls.

Track all of these. Different optimization strategies target different metrics.

For user-facing chat, TTFT is one important interaction metric; total completion time, output rate, interruptions, task success, and accessibility also matter. Set objectives from observed user behavior rather than assuming one latency metric dominates every interface.

For batch: total latency matters; throughput matters more.

For agents: tool call latency often dominates; optimizing the LLM doesn’t help if tools are slow.

Error observability

LLM-specific errors:

API errors. Rate limits, auth failures, server errors. Same as any API.

Validation errors. Structured output didn’t conform to schema. Track frequency by feature.

Content filter errors. Provider blocked the request. Track to detect prompt issues.

Tool errors. Specific tools failing. Track per tool.

Quality errors. Judge LLM scored output as bad. Track over time.

Hallucination signals. Detection of likely hallucinations (model claimed something not in the source). Hard to detect automatically but can be approximated.

Cost errors. Calls that cost much more than expected. Often indicate a bug.

Each gets its own dashboard. Each can trigger alerts.

A tagged interruption in a visible physical pipeline
AI-generated illustration of isolating one broken connection while keeping the surrounding application stages visible for diagnosis.

Debugging tools

When something breaks, you need to find it and understand it. The debugging surface includes:

Trace search. Find an event by trace ID, approved pseudonymous subject reference, or timestamp without exposing broader tenant data.

Call inspector. Show metadata by default. Reveal only approved, minimized request/response fields under role checks, access logging, purpose limitation, and retention rules; many deployments should never retain the full payload.

Trace timeline. For complex flows, see the chain of calls visually.

Controlled replay capability. Can an approved, minimized input be re-run in an isolated environment with tools disabled or sandboxed? Never replay production calls into live side effects, and do not assume retained payloads are justified merely because replay is useful.

Diff view. Compare two calls — different versions of the same prompt, different models — side by side.

Search by approved content or derived fields. Search for defined patterns without turning the telemetry system into an unrestricted corpus of customer prompts. Authorization, minimization, indexing, retention, and access logging apply to search as well as storage.

Products differ materially in these capabilities. Verify them in a representative trial and include integration, storage, privacy review, migration, and operating work in the comparison; list price alone does not prove buying is cheaper.

Privacy and PII

LLM observability logs are sensitive. Inputs may contain personal data; outputs may quote it. Raw payload capture is sometimes proposed for debugging, but it is not automatically necessary or lawful; define the purpose and consider synthetic reproduction, derived fields, or short-lived controlled samples first.

Practices:

Pseudonymization/tokenization. Replace direct identifiers with purpose-specific tokens and protect any re-identification lookup separately. If re-identification remains possible, the records are still personal data rather than anonymous “non-PII.”

Redaction at log time. Detect and redact PII before it lands in observability storage. Specific patterns (emails, phone numbers, SSNs) replaced with placeholders.

Tenant isolation. Multi-tenant observability data isolated per tenant. One tenant’s data not visible to another.

Access controls. Who can view raw inputs/outputs? Logged.

Retention policies. Logs older than X days are deleted or moved to cold storage. PII retention has legal limits in many jurisdictions.

Deletion and restriction workflow. Make telemetry records discoverable by the identifiers approved for that purpose and implement the action counsel defines across primary storage, indexes, exports, and backups. Do not translate the colloquial “right to be forgotten” into a blanket promise; GDPR erasure has conditions and exceptions.

The system needs controls proportionate to the personal data and purpose. The exact combination varies, but tenant isolation, authorized access, minimization, retention, deletion/restriction handling, and auditability must be decided before payload telemetry is enabled. The European Commission summarizes the conditions and exceptions for erasure requests.

Alerting

Thresholds and signals that warrant alerts:

Cost.

  • Spend or forecast exceeds the feature’s measured budget envelope.
  • Per-request cost exceeds a workload-specific ceiling.
  • Rate of change exceeds the baseline band for the same traffic mix.

Latency.

  • Tail latency breaches the product’s measured service objective.
  • Time to first token or completion breaches the interaction-specific objective.
  • Tool timeouts or queue time depart from the baseline band.

Error rate.

  • Error rate breaches the workload’s error budget.
  • A specific error type departs from its baseline band, including validation errors and rate limits.

Quality.

  • A calibrated metric changes beyond its repeated-run variance or a safety slice records any disallowed outcome.
  • User feedback negative rate > baseline.
  • Regeneration rate > baseline.

Pattern.

  • Specific bad phrases appearing more often.
  • Sudden change in input distribution.

Each alert should identify the feature, time window, observed and expected value, affected slice, trace links, and runbook. Do not diagnose a runaway in alert text unless loop or retry telemetry supports that diagnosis.

Multi-tenant considerations

For B2B SaaS apps:

Per-tenant metrics. Each customer can see their own usage, costs, and quality.

Per-tenant alerts. Specific to their thresholds.

Per-tenant debugging. Support can see a customer’s traces (with appropriate access controls).

Per-tenant configuration. Some customers may have different models, prompts, or policies. The observability layer reflects this.

For multi-tenant systems, tenant attribution and isolation may be necessary for support and billing, but raw trace access is not automatically justified. Give support the least data and capability needed, log access, and provide an escalation path for sensitive payloads.

Tool ecosystem (2026)

A landscape view of LLM observability tools as of writing:

Dedicated LLM observability:

  • Helicone, LangSmith, Phoenix (Arize), Braintrust, PromptLayer, and Weights & Biases Weave are candidates to verify against the requirements above.

General APM with LLM extensions:

  • Datadog LLM Observability and New Relic AI Monitoring are candidates when the organization already uses those platforms.
  • OpenTelemetry + your APM of choice. OTel has GenAI semantic conventions; instrument once, view in many tools.

Build-your-own:

  • A small database can be sufficient for a bounded system when it implements the required isolation, access, retention, deletion, and indexing controls.
  • Add a simple UI to search and display.
  • Integrate with your existing logging infrastructure.

Record the choice, rejected alternatives, data-flow review, exit/export test, and reassessment date. No default migration path fits every team.

A practical setup

Sequence by dependency and risk rather than a generic calendar:

Stage 1: Select a candidate against the requirements and run it on synthetic, non-sensitive traces. Verify export, access, redaction, retention, deletion, failure, and exit paths—not only ingestion.

Stage 2: Add dashboards for the defined service objectives, including per-feature cost, latency distributions, error classes, model/prompt versions, and missing-telemetry rates.

Stage 3: Set up workload-derived alerts on cost, loops, latency, and error classes.

Stage 4: Connect spans across multi-call flows and verify trace propagation through tools and queues.

Stage 5: Implement approved quality sampling and calibrate automated scores against human judgments.

Stage 6: Add user feedback only when its interpretation, privacy handling, and response workflow are defined.

Stage 7: Add tenant-specific views and debugging capabilities with authorization and access audits.

The acceptance record for each stage should include tests, data classifications, owners, failure behavior, and rollback. Skip a capability only with an explicit rationale and compensating control.

What goes wrong without it

A short catalog of incident patterns that recur at teams without proper LLM observability:

  • A feature retries calls in a tight loop and accumulates unexpected charges before a billing review.

  • A model update changes behavior, but requests do not record the model version, delaying attribution.

  • A prompt version regresses an important flow, but the deployment and response traces cannot be compared by version.

  • An agent loops, but missing step budgets and trace-level instrumentation obscure the repeated state.

  • A tool authentication failure triggers retries, but error-class and retry telemetry are not connected.

  • A prompt-injection path produces an unsafe disclosure, but missing data lineage and policy-decision telemetry prevent an impact assessment.

These are failure scenarios to test. Observability creates evidence for detection and investigation; it does not prevent the underlying failure unless enforcement controls act on the signals.

Instrument before the incident

LLM observability extends rather than replaces ordinary APM. Select the call, trace, quality, cost, latency, error, and debugging signals justified by the workload and threat model, and connect signals to tested response controls.

A production claim should instead show detection time by scenario, trace completeness, alert precision/recall where measurable, budget-control behavior, privacy tests, and restoration or rollback evidence. Instrument before launch, rehearse the runbooks, and publish only results you actually measured.

Read next

Continue through the same learning path with the next practical articles.