The n8n AI Agent node lets a model choose among configured tools. That flexibility adds non-determinism and a wider failure surface, so use an agent only after simpler deterministic routing is insufficient.
This article specifies a lead-triage design that receives a lead, retrieves permitted context, proposes scores and a draft, validates the result, and routes it to human review. It is not an importable workflow or a reported execution; the acceptance tests below define what remains before it is working end to end.
We assume you have n8n installed (self-hosted on a server or via n8n Cloud) and a working API key for Claude or OpenAI. If you’re new to n8n, work through their basic tutorial first.
The design also follows OWASP’s warning about excessive agency: minimize tool permissions, minimize autonomy, and require workflow-enforced approval for consequential actions. If lead data came from a third party or will be used for outreach, verify the source, notices, lawful basis, suppression, and channel rules before ingestion or contact; the European Commission explains that third-party data is not automatically reusable for marketing.
Do not let the first version auto-send replies to real leads. Route drafts to human review until you have logs, idempotency, score thresholds, and enough reviewed runs to know the workflow behaves under messy inputs.
What we’re building
The workflow:
- A new lead arrives via a webhook (from a form, an event, a CRM, etc.).
- The agent enriches the lead with company information (using web search).
- It scores the lead on three dimensions: fit, intent, urgency.
- It drafts a personalised response.
- Based on the score, it either:
- Queues a reply and CRM proposal for approval (for high-confidence high-fit leads),
- Drafts a reply for a human to review and sends a Slack notification (for medium leads),
- Or just logs and notifies without responding (for low-fit leads).
Parts of the pattern can generalise to low-consequence triage, but every new domain needs its own data, error, fairness, and professional-review analysis. Do not reuse a sales score design for employment, medical, legal, financial, child-safety, or construction decisions.
The companion JSON schema linked from this article defines the intake payload. Use it to validate the webhook input before the agent node sees it.
Add a pre-agent validation gate
The webhook should not pass arbitrary form data straight into the agent. Put a validation step between the trigger and the agent:
| Field | Rule | Failure behavior |
|---|---|---|
email | Required; trim and validate, normalize the domain conservatively, and preserve the local part unless the owning provider defines stronger canonicalization | Reject and notify owner |
message | Required, non-empty, maximum length | Reject or route to manual review |
source | Required enum such as website-form, event, crm | Reject unknown source |
timestamp | Required ISO timestamp or generated by webhook | Use receive time and flag |
lead_id | Required stable ID or generated idempotency key | Deduplicate before processing |
This gate protects the workflow from malformed submissions, duplicate webhook retries, and prompt-injection content hidden in form fields. The agent can still read the message, but the workflow decides whether the record is valid enough to process.
The mental model: agent = LLM + tools + loop
Before we build, the concept.
An “agent” in 2026 means an LLM that can use tools. Instead of producing a single response, the model decides which actions (called “tools”) to call. After each tool returns, the model sees the result and decides what to do next — another tool, another step, or a final answer.
n8n’s AI Agent node implements this loop. You give the model:
- A system prompt (its instructions and tone).
- A user prompt (the input for this run).
- A set of tools (other n8n nodes or sub-workflows the model can call).
The model decides which tools to call, in what order, with what arguments. After each tool returns, the model reconsiders. When it decides it has done enough, it returns a final answer.
This is fundamentally different from a static workflow because the order of steps is determined by the model, not by you. The skill of agent design is in:
- Giving the model the right tools (not too few, not too many).
- Writing a system prompt that scopes its behavior.
- Adding guardrails so it doesn’t go off-rails.
- Designing the output so downstream nodes can use it reliably.
Step 1: The trigger
Open n8n and create a new workflow. The trigger:
- Node: Webhook
- HTTP Method: POST
- Response Mode: “When last node finishes”
- Path: something like
/lead-triage
This webhook will receive lead submissions. n8n gives you a URL you can configure as the destination for your form submissions or your CRM’s outbound webhook.
For testing, save the workflow once so the webhook URL becomes active, and have a sample payload ready. A typical lead webhook payload might be:
{
"name": "Anna Lehtinen",
"email": "anna@somecompany.fi",
"company": "Some Company OÜ",
"role": "Head of Marketing",
"message": "Interested in your AI consulting services. We have a team of 10 and need help with prompt engineering training.",
"source": "website-form",
"timestamp": "2026-05-15T14:30:00Z"
}
Click “Test step” and submit the test payload to see the data flowing in.
Step 2: The AI Agent node
Add an AI Agent node after the webhook. Configure it:
- Agent / Tools Agent: Current n8n AI Agent nodes (1.82+) no longer expose an Agent Type dropdown; they run as a Tools Agent. Connect a chat model and the tools below. On older templates that still show Agent Type, choose Tools Agent — Conversational and other legacy types were removed.
- Chat Model: Claude (Anthropic) or OpenAI. Start with the current general-purpose model in your provider, then move to a faster or more capable tier only when your evals justify it. Reasoning modes can add latency and cost to every agent step.
- Memory: None for stateless triage (each lead is independent). For multi-turn agent conversations, use a memory node.
- System Message: This is where the agent’s behaviour lives. Use the template below.
- User Message: Pull the lead data from the webhook.
The system message:
You are a lead-triage agent for [Your Company Name], an AI consulting firm.
Your job is to process incoming leads and produce a structured triage decision.
For each lead, you must:
1. Use the `enrich_lead` tool to gather context about the company.
2. Score the lead on three dimensions:
- Fit: does the lead match our ideal customer profile?
- Companies of 10-200 people in B2B, manufacturing, or professional services.
- Roles in marketing, operations, engineering leadership, or executive.
- Intent: how serious is the inquiry?
- "Just curious" vs "actively evaluating" vs "ready to buy."
- Urgency: is there a stated or implied timeline?
3. Use the `score_lead` tool to record the scores.
4. Use the `draft_response` tool to produce a personalised reply.
5. Use the `propose_route` tool with one of: "review_priority", "human_review", "log_only".
Routing rules (evaluate in this order; first match wins):
- "review_priority" if Fit >= 7/10 AND Intent >= 7/10. The reply receives priority human review.
- "human_review" if Fit >= 5/10 OR Intent >= 5/10. A human will check before sending.
- "log_only" otherwise (Fit < 5/10 AND Intent < 5/10). We just track and move on.
Never invent information. If something is unclear, mark [unclear] in your scoring rationale.
Always end by returning a JSON object with:
{
"fit_score": <1-10>,
"intent_score": <1-10>,
"urgency_score": <1-10>,
"reasoning": "<2-3 sentences>",
"drafted_response": "<the email body>",
"routing": "<review_priority|human_review|log_only>"
}
Notice the structure. We have:
- A clear job statement.
- An explicit process (steps 1-5).
- Explicit scoring criteria.
- Explicit routing logic.
- A required output format.
The agent will not always follow this perfectly. But the more concrete the system message, the more reliably it executes the same shape every time.
The JSON object is not enough by itself. Add a validation step after the agent node and reject runs where scores are missing, routing is outside the allowed enum, or the drafted response is empty.
Step 3: The tools
The agent needs tools to call. In n8n, tools are configured under the agent node and can be:
- Sub-workflows.
- HTTP requests.
- Built-in tool nodes.
Let’s build four tools for our agent.
Tool 1: enrich_lead
A sub-workflow that:
- Takes a company name and email domain as input.
- Uses an approved search or company-data API whose terms permit the use.
- Returns structured claims with canonical source URLs and retrieval dates; unknown size or identity remains unknown.
The tool’s description (which the agent reads to decide when to call it):
Retrieves permitted public context for a lead company. Input: company name and email domain. Output: verified claims with source URL and retrieval date, plus ambiguity/errors. Do not infer identity, size, or news without a source.
Tool 2: score_lead
A deterministic validation tool that:
- Accepts proposed scores and checks type, range, required rationale, and allowed labels.
- Returns validation errors or a normalised score object.
- Has no database, sheet, CRM, email, or other write credential.
Persist only after the agent’s final output passes the same server-side schema.
The tool’s description:
Validates proposed lead scores without persisting them. Input: fit_score, intent_score, urgency_score, and rationale. Output:
{valid, errors, normalized_scores}. This tool cannot write records or send messages.
Tool 3: draft_response
A sub-workflow that takes the lead context and the scores and produces a personalised email draft. This sub-workflow internally calls another AI node with a specific drafting prompt:
Draft a personalised response to a B2B inquiry. Inputs: the original lead message, the company enrichment summary, the fit/intent/urgency scores.
Voice: warm, direct, no corporate filler. Acknowledge the specific request. Reference enrichment only when the claim has a source and is relevant; otherwise omit it. End with a proposed next step for human review.
Length: 80-120 words.
The tool’s description:
Drafts a personalised email reply to the lead. Input: the lead message, enrichment summary, and scores. Output: an email draft.
Tool 4: propose_route
This tool records one of three proposed routes; it does not send email or write to the CRM:
review_priority: place the draft in the priority human-review queue.human_review: place the draft in the standard human-review queue.log_only: record the triage outcome without preparing an outbound action.
A deterministic Switch node after schema validation enforces the allowed enum and sends review_priority and human_review to an approval queue. Only a separate, approval-gated sub-workflow owns customer-visible writes.
Implement this as a side-effect-free sub-workflow that returns a proposal object. After the Agent node finishes, a deterministic schema-validation node and Switch node decide which branch may persist the proposal. None of those branches may reach a send node without the separate approval workflow.
The tool’s description:
Proposes a route. Input: routing decision (
review_priority,human_review, orlog_only), scores, rationale, and draft. Output: an in-memory proposal object for deterministic schema validation. This tool cannot persist, send email, or update the CRM.
Step 4: Test the agent
With the agent configured and the four tools attached, run a test using the sample payload.
What you should see in n8n’s execution view:
- Webhook receives the payload.
- AI Agent starts.
- Agent calls
enrich_lead— you see the tool execute and return. - Agent selects the next step (you may or may not see intermediate traces depending on model and settings).
- Agent calls
score_lead. - Agent calls
draft_response. - Agent calls
propose_routewith one of the three routing options. - Agent returns final JSON.
If something goes wrong, n8n’s debug panel shows you the messages between the agent and its tools. The most common issues:
- Tool description not specific enough. The model cannot infer when the tool is appropriate. Make descriptions more concrete.
- Tool input/output schema mismatched. The agent can’t pass the right arguments. Be explicit about the schema.
- Agent loops forever. It keeps calling tools without resolving. Add a max-iterations limit and re-examine your system prompt.

Step 5: Add guardrails
A bare agent is unsafe in production. Six guardrails to add before you trust it with real traffic:
1. Max iterations. Set a finite limit from the smallest number your successful eval cases require. Test that hitting the limit routes to a human and cannot leave a partial side effect.
2. Approval gate for every proposed reply. review_priority changes queue order; it does not authorise send. Keep customer communication behind an authenticated human approval until a separately approved policy says otherwise, and never infer safety from elapsed weeks.
3. Allowlist for outbound actions. Configure your CRM tool and email tool to only act on records that match your expected pattern. Prevents the agent from sending email to the wrong address or creating records for non-leads.
4. Logging. Emit approved metadata for every run: stable run/lead reference, workflow and model versions, tool names and outcomes, validation result, route, approver, retries, and errors. Raw lead input, enrichment results, drafts, and tool arguments contain personal or confidential data and require a separate purpose, redaction, access, and retention decision.
5. Cost limits. Set finite agent iterations and workflow timeouts, and configure spend/rate alerts or limits with the model provider. n8n plan limits are described in executions and feature terms; do not assume n8n Cloud enforces a daily budget for a bring-your-own provider key. Track provider usage and test the kill switch.
6. Decision ownership. The model can recommend review_priority, human_review, or log_only; the workflow enforces the final rule. Keep the routing enum, score thresholds, and approval requirements outside the prompt so they are testable and visible.
Step 6: Production hardening
A few patterns that turn a working prototype into something you can trust:
Idempotency. Make sure that if the same lead is processed twice (because of a webhook retry or a manual re-run), it cannot create duplicate records or messages. A read-then-write check is race-prone: atomically claim a unique key in a database and use the same key on every downstream write. Follow the atomic claim, lease, approval-token, and outbox design.
Error handling. Wrap each tool call in error handling. If enrichment is unavailable or company identity is ambiguous, the workflow should flag the missing data and route to human review; it must not invent personalised facts merely to complete the draft.
Observability. Track key metrics: average run time, tool call frequency, percentage routed to each path. Anomalies are signals.
Pilot review. Review every run in a consented, bounded pilot against the source lead. Size the pilot to cover source types, languages, missing fields, ambiguous companies, injection attempts, and priority classes. Track wrong enrichment, score, route, deadline, and draft. A fixed count of 50 does not establish a reliability level.
An enforceable kill switch. Gate execution and every side-effect dispatcher outside the model so an authorized operator can pause new and resumed runs without redeploying. Test that the disabled state blocks queued and in-flight sends; a prompt instruction or value that only the agent “checks” is not a kill switch.
The most important design decision: what tools to give the agent
The biggest determinant of agent quality is the tool set. Two failure modes:
Too few tools. The agent can’t do its job. It tries to fake its way around missing capabilities, often by hallucinating.
Too many tools. The agent gets confused, picks the wrong tool, or wastes iterations exploring. Quality degrades.
A good rule: start with the minimum viable tool set, and only add tools when the agent demonstrably needs them.
For lead triage, the four tools we chose are roughly right. You might add:
- A “lookup_existing_customer” tool to check if the lead is already a customer.
- A “schedule_meeting” tool that integrates with your calendar.
- A “translate” tool if leads come in multiple languages.
But every new tool is a new decision the agent has to make. Each one should genuinely earn its place.
Patterns that generalise
The same control ideas may help other triage workflows, but the labels and actions below are not transferable without domain review:
Support ticket triage. Retrieve permitted customer history and propose category/priority; keep account, safety, refund, entitlement, and customer-message actions behind policy and human gates.
Employment workflows. Do not adapt the lead score into interview/reject automation. Employment decisions can create legal and discrimination risk and require qualified HR/legal review, accessibility controls, bias evaluation, worker/applicant transparency, and meaningful human decision-making.
Press inquiry handling. Replace enrichment with “lookup_publication,” and routing with priority-based response.
Customer feedback routing. Replace enrichment with sentiment analysis and product categorization.
Procurement requests. Replace enrichment with vendor lookups, scoring with policy compliance, and routing with approval flows.
A reusable low-consequence shape is: validate an inbound event → retrieve minimum permitted context → ask for a structured proposal → validate deterministically → let authorised policy/humans decide → execute through idempotent gated tools. The model proposes; it does not own the consequential decision.
When NOT to use an agent
Some workflows don’t benefit from an agent. If the logic is fully deterministic — “always do A, then B, then C” — a regular n8n workflow without an agent is faster, cheaper, and more reliable.
The agent earns its place when:
- The number of possible paths is large.
- The right path depends on judgement, not strict rules.
- Some decisions require synthesising information across multiple sources.
If your decision tree is a few if-then-else statements, just use if-then-else nodes. Save the agent for cases where the if-then-else gets unmanageable.
Build it once on real work
An AI agent in n8n is a workflow where a model may choose among configured tools. A production design constrains that choice and keeps accountable humans and deterministic policy in control of consequential decisions.
The design is complete only after a duplicate-delivery test, invalid-output test, provider-timeout test, rejected-approval test, and a check that the send node is unreachable without approval. Measure build time, latency, correction rate, and cost on your own instance; this article does not promise a setup time or production outcome.
Build and test it first with synthetic or consented non-production records. Reuse the control pattern — narrow tools, schemas, idempotency, stop rules, and review gates — while redoing the domain and risk analysis for each new workflow.



