Design an AI customer support agent: triage, knowledge, actions, and escalation
Intermediate11 min readAutomations

Design an AI customer support agent: triage, knowledge, actions, and escalation

A queue-evaluated reference design for support triage, retrieval, response drafting, controlled actions, and human escalation, with the policy and measurement boundaries needed for a safe pilot.

What you should be able to do

A support agent is only as useful as its knowledge, tools, escalation policy, and measured results. Start with a narrow ticket class, keep a human path, and expand only when resolution and recontact data justify it.

Saved only in this browser.
In this article

Headline claims about ticket resolution, savings, and response time may describe a particular vendor deployment, but they do not establish what your queue can automate. The result depends on scope, policy, knowledge quality, tool access, escalation, and how “resolution” is measured.

There is no defensible universal automation rate for a “typical SaaS queue.” Password resets, billing disputes, outages, and product bugs have different handling limits, and vendors count “resolution” differently. An agent can produce unsupported or policy-violating responses even when its language sounds confident. The architecture and the measurement design matter more than a headline percentage.

This article presents a reference design to test against a labelled sample of your own queue. Its categories, time windows, thresholds, prompts, and workflow stages are examples, not defaults or performance claims. Keep only the parts that pass your policy, security review, and evaluation criteria.

The four jobs of a support agent

A useful AI support agent does four things, in order:

  1. Understand the ticket. What is the customer actually asking? What are they emotionally bringing to this? What category of issue is this?
  2. Look up the right context. The customer’s account, their history with you, the relevant documentation, similar resolved tickets.
  3. Decide what to do. Reply with the answer, ask a clarifying question, route to a human, take an action on the account.
  4. Execute an authorised decision. Send the reply, ask the question, escalate, or perform an approved account action, while recording the evidence, authorisation, and outcome needed for operations and audit.

Many support-agent failures originate in missing context or unclear routing, but model capability and evaluation still matter. Treat the system as a whole.

The architecture

One reference workflow is:

Incoming ticket
    ↓
[Triage agent: classify, prioritise, route]
    ↓
[Context gathering: customer data, history, knowledge base RAG]
    ↓
[Decision agent: propose response or action]
    ↓
[Policy gate: authorise, confirm, block, or route]
    ↓
[Response drafter / controlled action executor]
    ↓
[Response quality check, send, or escalate]

The boxes are responsibilities, not a required number of models or services. They can be combined or separated in n8n, an agent framework, or an owned service, provided the authorisation boundary remains outside the model. The flow resembles the routing pattern in Anthropic’s guide to building effective agents, which includes customer-service routing as an example; it is not a platform-independent guarantee.

We will walk through each step.

Step 1: Triage

The triage agent receives the raw incoming ticket and classifies it.

An illustrative triage prompt follows. Adapt the labels and thresholds to your queue, then test category-specific false positives and false negatives before routing live tickets:

You are a triage agent for [Company]'s customer support. Classify each incoming ticket on three dimensions:

1. CATEGORY: one of
   - account_access (login, password, MFA, account locked)
   - billing (charges, refunds, plan changes, invoices)
   - product_question (how-to, feature questions, configuration)
   - bug_report (something broken or unexpected)
   - feature_request (asking for something we don't have)
   - complaint (frustrated customer, not a specific technical issue)
   - other

2. URGENCY: one of "critical" (production down, billing dispute), "normal", "low" (informational).

3. EMOTIONAL_TONE: one of "calm", "frustrated", "very_angry". Be honest.

Output JSON with `category`, `urgency`, `emotional_tone`, `confidence`, and `needs_review`. Set `needs_review` when the evidence is insufficient or the result fails the validated threshold for that category.

Run triage on the least expensive, lowest-latency model that passes your classification and routing evaluations. A more capable model may still be necessary for ambiguous or multilingual queues; decide from measured errors, not a fixed model label.

The output of triage can feed two policy decisions:

  • In one policy, high-urgency or very-angry tickets go directly to a human. Another queue may use different signals or require deterministic incident rules.
  • Category can restrict which knowledge source and tools are available downstream; it must not grant permissions by itself.

Step 2: Context gathering

Context quality is a major variable in support quality. Without relevant, authorised evidence, the model may fill gaps with unsupported inferences.

Three possible sources of context are:

Customer data. Who is this customer? Plan, account age, recent activity, payment status, any open issues. This usually comes from your CRM or product database via an API call.

Conversation history. Has this customer contacted you before? What about? How was it resolved? Avoid the “I just told you that yesterday” failure mode.

Knowledge base. Documentation, help-center articles, and internal runbooks. Retrieval may use filters, keyword search, dense retrieval, hybrid search, reranking, or a product-specific combination. Choose the method and result count from retrieval evaluations, not from the generic label “RAG.” (See build a personal RAG and production RAG.)

The following windows and result counts are illustrative placeholders. Set them from queue history, privacy and retention rules, latency budgets, and retrieval evaluations:

Given the ticket [content], gather context:

1. Look up the customer by email. If found, retrieve plan, account_age_days, recent_actions (last 7 days), open_tickets.

2. Look up the customer's ticket history (last 90 days). Retrieve up to 5 most recent tickets with their resolution.

3. Search the knowledge base for relevant articles. Retrieve top 3 by semantic similarity. Include article titles, summaries, and URLs.

4. Search resolved tickets in our database for similar issues. Retrieve top 2 with their resolutions.

Combine into a context object.

This step gives the agent evidence to work with, but its coverage and latency depend on the connected systems, retrieval design, and verified service-level targets. Record document identifiers and versions so a reviewer can reconstruct which evidence was available.

Privacy boundary. Customer context is permissioned data, not a prompt convenience. Retrieve only the fields needed for the ticket, enforce tenant and role access before retrieval, redact secrets and unnecessary personal data, and apply approved retention rules to prompts, tool results, traces, and drafts. Never rely on the model to decide which records it was authorised to see.

Step 3: The reasoning agent

Now the agent proposes what to do. This prompt is an example decision policy, not authorisation to execute an action. Replace its fixed money, confidence, and emotion thresholds with values approved for each ticket category and jurisdiction:

You are a customer support specialist for [Company]. Your job is to resolve the customer's issue.

For each ticket:

1. Read the ticket and the context carefully. The context includes the customer's account, their history with us, and relevant documentation.

2. Decide on one of these actions:
   - RESOLVE: you have a confident answer or solution. Draft a response.
   - CLARIFY: you need more information. Draft a clarifying question.
   - ESCALATE: this needs a human. Explain why.
   - ACT_AND_RESOLVE: propose an account action (issue refund, reset password, change plan, etc.) and draft the response that would follow. Do not execute it; the downstream policy gate decides whether it may run.

3. Your tone is direct, warm, and competent. Match the customer's register. Never patronise. Never apologise more than once. Never use "we appreciate your patience."

4. When citing documentation, link to the specific article. Do not paraphrase from memory.

5. If the customer is frustrated, acknowledge it briefly and clearly, then move to the resolution.

6. Always escalate if:
   - The customer asks to speak to a human.
   - The issue involves a financial dispute over €100 / $100.
   - The case fails the validated confidence or policy threshold for its category.
   - The customer's tone is angry and the issue is not a simple one-step resolution.
   - The issue involves a security or privacy concern.
   - The issue involves a complaint about a person on our team.

7. Your output must be JSON:
{
  "action": "<resolve|clarify|escalate|act_and_resolve>",
  "confidence": <0.0-1.0>,
  "reasoning": "<brief explanation>",
  "response_draft": "<the email body>",
  "escalation_reason": "<if applicable>",
  "action_to_take": "<if act_and_resolve, the specific action and arguments>"
}

This is the heart of the agent. Choose the least expensive model that meets your disposition, answer-quality, safety, and escalation thresholds on representative tickets. Re-evaluate it when the model, prompt, tools, or queue changes. The reasoning field in this example should contain a short evidence-and-policy rationale for an operator, not private chain-of-thought and not proof that an action was authorised.

Step 4: Action execution

For RESOLVE and CLARIFY, the proposed response still passes the response-quality and policy checks before it is sent.

For ESCALATE, route the ticket to a human queue with the customer message, retrieved evidence, relevant policy rule, attempted steps, and reason for escalation. Do not substitute hidden model reasoning for that operator-facing record.

For ACT_AND_RESOLVE, the model proposes an action. A separate control path decides whether it may run. OWASP’s excessive-agency guidance recommends minimum functionality and permissions, execution in the user’s security context, downstream authorisation, and human approval for high-impact actions. Apply those controls to the support workflow:

  • Least-privilege allowlist. Expose only narrowly scoped operations. A policy might permit refunds up to an illustrative €50 while requiring review above that amount, but the real boundary must come from the authorised policy, not from this article or the prompt.
  • Identity and authorisation. Resolve the authenticated customer, tenant, operator, and granted scope before execution. Enforce access in the downstream service on every call; a model’s classification or confidence cannot grant access.
  • Validated arguments. Constrain action names and arguments with a schema, reject unexpected fields, and re-check account, currency, amount, destination, and policy immediately before the side effect. Where the provider supports schema-constrained tool calls, enable it; OpenAI’s strict function-calling mode, for example, enforces schema adherence but does not establish authorisation or factual correctness.
  • Confirmation and approval. Require explicit customer confirmation or human approval where risk, policy, law, or ambiguity warrants it. Security-sensitive recovery must follow the verified identity-recovery process rather than a generic reset action.
  • Idempotency and concurrency. Give each action an idempotency key, prevent duplicate execution on retries, and handle stale account state or competing updates.
  • Reversibility and failure handling. Prefer staged or reversible operations, define rollback or reconciliation for partial failure, and route uncertain outcomes to a human instead of retrying blindly.
  • Security controls. Apply rate limits, tool timeouts, tenant isolation, secret handling, and abuse monitoring independently of the model.
  • Audit record. Log the request identifier, actor and customer identity, redacted inputs, evidence identifiers and versions, policy and authorisation result, confirmation or approver, exact action arguments, tool result, and rollback or error state. A model-generated rationale may help review, but it is not the audit trail.

Step 5: Quality check

Use two different controls. Before any side effect, deterministic policy and authorisation checks must block, approve, or route the proposed action. Separately, a model or rules-based response reviewer can detect answer-quality problems before a message is sent. Treat that reviewer as an evaluated detector with known false-positive and false-negative rates, not as an infallible judge or an authorisation service.

You are a quality reviewer for AI-generated customer support responses.

Given the original ticket and the drafted response, check:

1. Does the response actually address the customer's question?
2. Is it accurate based on the context provided (no hallucinated facts)?
3. Is the tone right (warm, direct, not patronising, not over-apologetic)?
4. Are any links broken or wrong?
5. Does it contain any of these red flags:
   - Promising something we cannot deliver
   - Apologising for things that aren't our fault
   - Sounding angry or sarcastic
   - Using internal jargon
   - Disclosing internal information

Output: APPROVE or REVISE (with specific suggested fixes).

If the quality check returns APPROVE and the response passes policy, it can be sent. If it returns REVISE, apply a bounded revision and re-check it, or route it to a human. Put a retry limit on automated revision so a failed detector does not create a loop.

Whether this quality gate is worth its cost is an empirical question. Log how often it changes the disposition, how many bad replies it catches, and how many good replies it blocks; retain it only if those measurements justify the extra latency and model call.

Make the knowledge base testable

The knowledge base is one critical factor in agent quality. If your help center is stale, contradictory, or incomplete, the agent can produce confident but unsupported answers.

Practical principles:

Audit before deploying. Sample the highest-volume and highest-risk ticket types and verify that the knowledge base has the right answer for each. Fill gaps, resolve contradictions, and update stale articles. Expand the sample until your acceptance criteria are supported by evidence from your queue.

Structure for the selected retrieval method. Focused sections, clear titles, stable identifiers, and explicit applicability can help retrieval, but chunking and article length are implementation choices. Test whether the required passage and its scope are retrieved for representative questions.

Include explicit “do NOT” sections. Many support tickets are about how to do something the customer should not do. Knowledge base articles should explicitly say “if you are trying to X, here is why we don’t recommend that, and here is the alternative.”

Represent applicability. Metadata such as “Free plan only,” “EU customers only,” or “iOS app only” can support filters. Enforce these restrictions in the retrieval layer and test that conflicting or out-of-scope material is excluded.

Review on ownership and change triggers. Give each knowledge area an owner and a review interval appropriate to its risk and change rate. Re-review affected material when products, policies, incidents, or regulations change.

Question cards are paired with source pages on a wooden table
AI-generated illustration of testing a customer-support knowledge base against representative questions.

Calibrate escalation by policy and evaluation

A system that escalates too broadly adds queue load; one that escalates too narrowly can create customer and security harm. Define rules by category, consequence, evidence quality, customer choice, and measured error rates. A starting policy might include:

Route to a human or specialist queue:

  • Explicit human requests
  • Anger above a threshold (especially after one bad agent turn)
  • Disputes involving real money
  • Security or privacy concerns
  • Health, safety, or legal implications
  • Repeated tickets from the same customer about the same issue
  • Cases that fail the validated confidence or policy threshold for their category

Candidates for automation after they pass the relevant evaluations:

  • Trivial questions with clear answers in the KB
  • Account housekeeping (password reset, basic profile changes)
  • Status queries (“did my refund go through?”)
  • Feature requests (route to product team, not human support)

These are not universal lists. A password reset, refund status, or account change can be high risk in one product and routine in another. Automate only when the response or action is within policy, the caller’s identity and authority are verified, the tool is tightly scoped, and the case passes the category’s measured acceptance rule.

The middle ground is where the agent’s judgement matters. Build instrumentation that lets you see: of all the cases the agent could have escalated but didn’t, what fraction did the customer come back about? Of all the cases the agent escalated, how many did the human resolve trivially?

Derive an automation target from your own queue

Do not start with a vendor’s resolution-rate target. Label a representative sample of your own recent queue into categories such as:

  • simple, clearly documented questions;
  • questions that need account context or a controlled tool;
  • complex troubleshooting, emotional situations, or policy decisions;
  • bug reports and feature requests that belong with product or engineering.

For each category, test whether the agent can produce a correct disposition and response under your actual policy. The sum of categories that clear your acceptance thresholds is your initial automation ceiling. Recalculate it after knowledge-base, tool, or policy changes; do not back-solve the labels to reach a promised percentage.

Measure customer outcomes

Measure what customers value in your own queue. Common candidate metrics are:

  • time to a correct resolution;
  • answer accuracy and policy compliance;
  • recontact rate, customer effort, and satisfaction;
  • ability to reach a human when the automation fails.

Do not infer preference from speed alone. A fast wrong answer, or a bot that hides the human path, can make the experience worse than a queue with an explicit wait time.

A repeated automated loop without a usable human path is a foreseeable failure mode. Measure repeated contacts and abandoned sessions, cap automated retries, and make escalation discoverable.

A few specific patterns

Personalisation can be relevant. “Hi Anna, I see you’re on our Pro plan and have been with us since 2023” lands differently from “Hi Customer.” Use only approved details that help resolve the case, and avoid details that feel surveillant.

Acknowledge a verified wait when it is relevant. Use the ticket system’s timestamps and your actual service-level policy rather than inventing how long the customer waited or implying that a target was missed.

Confirm the relevant detail. “You mentioned that your import was failing on records with special characters in the company name.” Use this only when it accurately reflects the ticket; repetition is not evidence that the system understood it.

End with a verified next step. “The refund was accepted by [payment system] at [time]. Its current settlement window is [verified policy or provider window].” Do not claim an action succeeded or invent a delivery window from the model’s draft.

Don’t apologise unprompted. “I’m so sorry for the inconvenience” before knowing what happened comes across as insincere. Apologise once, specifically, where warranted.

A worked example

The customer writes:

Hi, I’ve been trying to log in for three days and it just keeps saying my password is wrong. I’m sure it’s the right password — I’ve used it for two years. I’m starting to think you’ve been hacked.

A safer draft, after the system has verified the account and the approved recovery path, is:

Hi Anna,

I understand why three days of failed sign-ins would be worrying. The sign-in records available to support show repeated failed attempts, but they do not establish who made them or whether your account was accessed. I have not changed your password or multi-factor settings.

Please use the account-recovery link on our verified sign-in page: [approved recovery URL]. Before any reset is completed, the recovery flow will verify your identity. Do not share a password, one-time code, recovery code, or reset link with support.

If you do not recognise the attempts, cannot complete the verified recovery flow, or notice an unfamiliar session or account change, reply here and I will route the case to our account-security team. Their response target is [verified security-queue SLA].

— AI Expert Support

That draft distinguishes observed evidence from inference, avoids claiming that there was no compromise, does not expose or select a recovery address, and makes the security escalation and response-time statement explicit placeholders. The final version still needs the company’s verified URL, identity process, and current service-level target.

What to build first

Set the resolution target from a labelled baseline and pilot data, not from this article or a vendor case study. The model is only one part of the system. Four major levers are:

  1. A governed, evaluated knowledge base.
  2. Solid context-gathering (customer data, history, KB retrieval, similar resolved tickets).
  3. A reasoning agent with clear decision criteria and escalation rules.
  4. Control gates (authorisation, allowlists, confirmation, response checks, and audit logging).

These controls can support a useful pilot, but they do not guarantee better support or lower human workload.

Start with a narrow, reversible pilot. Measure correct disposition, grounded-answer accuracy, policy compliance, unauthorised-action attempts, duplicate or failed actions, recontact rate, time to correct resolution, customer effort, escalation precision and recall, and the human queue load created by false positives. Segment the results by ticket category, language, customer group where appropriate, and tool action so an aggregate score does not hide a dangerous slice.

Residual risk remains after the pilot: retrieval can omit or surface stale evidence, identity signals can be wrong, policy may be incomplete, reviewers can miss unsafe drafts, integrations can fail between approval and execution, and customers can misunderstand an automated response. Keep a visible human path, incident stop control, monitored rollback or reconciliation process, and named owners for policy, knowledge, tools, and evaluation. Expand the scope only when the measured benefits and residual risk support the next category.

Read next

Continue through the same learning path with the next practical articles.