MCP (Model Context Protocol) is implemented by multiple agent clients and development tools. That adoption makes it a useful integration boundary; it does not make any server portable, secure, or production-ready without client and threat-model testing.
If you want LLM agents to interact with your service, an MCP server is how. And once you’ve built one or two, you’ll see that the protocol itself is small. The interesting engineering is in everything around it: schema design, error handling, auth, streaming, performance, observability.
This article has two deliberately separate outputs: a minimal stdio server you can run, and a production design checklist. It does not pretend that the fragments form a deployed authenticated service. The minimal code targets @modelcontextprotocol/sdk@1.30.0; the current v2 SDK uses split packages (@modelcontextprotocol/server, @modelcontextprotocol/node, and framework adapters), so follow the official server guide when starting new v2 work. Do not mix v1 and v2 imports.
What MCP is, briefly
MCP is a client-server protocol where:
- Servers expose tools, resources, and prompts.
- Clients are typically LLM agents that consume them.
The protocol uses JSON-RPC 2.0; read the official specification for the pinned revision. Transports are stdio for local processes and Streamable HTTP for remote servers—the older HTTP+SSE transport was deprecated in the 2025-03-26 revision. The HTTP authorization specification is based on OAuth; application gateways may add other credential schemes, but an API key is not a substitute for claiming MCP authorization-spec conformance.
The server’s job: expose useful capabilities to LLMs in a way they can discover and use.
The basic structure
Create a clean directory and pin the same major version as this example:
npm init -y
npm install @modelcontextprotocol/sdk@1.30.0 zod@3
npm install --save-dev typescript@5 @types/node
Then add this minimal server using the v1 high-level McpServer API:
import { McpServer } from "@modelcontextprotocol/sdk/server/mcp.js";
import { StdioServerTransport } from "@modelcontextprotocol/sdk/server/stdio.js";
import { z } from "zod";
const server = new McpServer({
name: "my-server",
version: "1.0.0",
});
server.registerTool(
"echo",
{
description: "Echo back the provided text.",
inputSchema: { text: z.string() },
},
async ({ text }) => ({
content: [{ type: "text", text }],
})
);
const transport = new StdioServerTransport();
await server.connect(transport);
(The same SDK also exposes the lower-level Server class plus setRequestHandler for ListToolsRequestSchema / CallToolRequestSchema if you want full control of the request handlers — but for most servers McpServer.registerTool is shorter and harder to get wrong.)
That’s the skeleton. What you put into the tool handlers — and how — is where the work is.
Pattern 1: Tool design philosophy
The first decision: what tools do you expose, at what granularity?
A common failure is mechanically exposing every underlying endpoint as a tool. A large irrelevant tool set increases schema context and can increase selection errors for some models and tasks. Compare task-focused tools with the endpoint-shaped baseline and expose only the authorized subset needed for the session.
Better: design tools for the way agents want to use them. Each tool does one well-defined thing, takes well-defined inputs, returns well-defined outputs.
A few principles:
One concept per tool. Don’t have a manage_customer tool that does 12 different things. Have search_customers, get_customer, update_customer_email, archive_customer — each focused.
Right granularity. Too granular and the agent needs many calls; too coarse and it can’t precisely do what’s needed. Aim for “operations a human would name.”
Action verbs. search_documents, not documents. Tools should be named by what they do.
Read-vs-write distinction. Read tools are safer; write tools have side effects. Distinguish in naming (list_x vs create_x) and treat differently (require explicit confirmation, idempotency keys, etc.).
Aggregate when useful. A get_customer_profile that returns customer + recent orders + support tickets in one call is often better than three separate calls. The agent gets context in one shot.
There is no universal correct tool count. Expose only the tools relevant and authorized for the current task, then measure selection errors as you add or remove tools.
Pattern 2: Schema design
Every tool has an input schema (parameters the LLM must provide) and an output (what your tool returns). Schemas are not just for validation; they’re prompt engineering.
Using Zod for input schemas:
const searchCustomersSchema = z.object({
query: z.string().describe(
"Search term: name, email, or company. Be specific to avoid too many matches."
),
limit: z.number().int().min(1).max(50).default(10).describe(
"Maximum results to return. Default 10, max 50."
),
filters: z.object({
tier: z.enum(["free", "pro", "enterprise"]).optional().describe(
"Filter to specific customer tier"
),
status: z.enum(["active", "trial", "churned"]).optional().describe(
"Filter by customer status"
),
}).optional(),
});
Notice:
- Every field has a
.describe(). The description is what the LLM reads. - Enums are explicit. Free-form strings are restricted where possible.
- Defaults are sensible.
- Constraints (min/max, length) are explicit.
- Optional vs required is clear.
The descriptions matter enormously. “search term” is unhelpful; “Search term: name, email, or company. Be specific to avoid too many matches” is useful guidance to the LLM.
Pattern 3: Output shape
Output is what the LLM sees and acts on. Good output design dramatically improves LLM behavior.
Structured outputs.
type SearchResult = {
customers: Customer[];
total_matches: number;
truncated: boolean;
next_page_cursor?: string;
};
With context.
{
customers: [...],
total_matches: 47,
truncated: true,
next_page_cursor: "abc",
message: "Found 47 matches; showing first 10. Use next_page_cursor to get more."
}
The message field is human-readable guidance. LLMs use it.
With errors handled gracefully.
{
error: "ambiguous_query",
message: "Search term 'john' matched 247 customers. Please be more specific.",
suggestion: "Try including a company name or email domain.",
partial_results: [...] // top 3 by relevance, optional
}
The error is structured (machine-readable), but also includes a message and a suggestion (LLM-readable). The LLM can adapt — either ask the user for clarification or refine the query.
Sized appropriately.
A tool that returns an unbounded record set can exceed context, cost, latency, and data-exposure limits. Enforce a maximum, paginate with stable cursors, and return only fields authorized and needed for the task. Summaries are derived data and need provenance when the exact records matter.
Pattern 4: Error semantics
Tools fail. How they communicate failure to the LLM determines whether the LLM recovers gracefully or compounds the error.
Categories of error.
type ToolError =
| { type: "validation"; message: string; field?: string }
| { type: "auth"; message: string }
| { type: "not_found"; message: string; suggestion?: string }
| { type: "conflict"; message: string; resolution?: string }
| { type: "rate_limit"; message: string; retry_after_seconds: number }
| { type: "service_unavailable"; message: string; retryable: boolean }
| { type: "internal"; message: string; trace_id: string };
Each category has different semantics. The LLM should respond differently:
validation: fix the input and retry.not_found: tell the user, or try a different search.conflict: ask for resolution.rate_limit: wait and retry.service_unavailable: try fallback or notify user.internal: give up, surface to user.
Documenting these in your server makes the LLM more capable.
Error formatting.
Return errors as structured data, with clear, actionable messages:
{
error: {
type: "validation",
message: "The email address is not in a valid format.",
field: "email",
suggestion: "Provide a valid email address like 'name@example.com'."
}
}
Avoid:
{
error: "Invalid input"
}
The first lets the LLM recover. The second leaves it guessing.
Pattern 5: Auth and authorization
Remote production MCP servers normally need authenticated, authorized callers, and local stdio servers still inherit the permissions of the launching process. Network reachability alone must never grant tool access.
Authentication: who is calling?
Common patterns:
- API key. Simple, common, works for service-to-service. Issue per consumer; rotate periodically.
- OAuth. For multi-user systems where end-users authorize agents. More complex but the right answer for many use cases.
- mTLS. For high-security environments. Mutual TLS certificates for both sides.
Implementation depends on transport. Over HTTP, validate bearer credentials before dispatching the MCP request. In v1, the official bearer-auth middleware attaches validated authInfo to the handler’s extra data; in v2, use the current resource-server and request-state APIs. Do not invent an application-only context.caller field.
server.registerTool("who_am_i", {
description: "Return the authenticated caller identity.",
inputSchema: {},
}, async (_input, extra) => {
if (!extra.authInfo) {
return { isError: true, content: [{ type: "text", text: "Authentication required" }] };
}
const subject = String(extra.authInfo.extra?.sub ?? extra.authInfo.clientId);
return { content: [{ type: "text", text: JSON.stringify({ subject }) }] };
});
This handler assumes the HTTP transport has already run the official authentication middleware; registering it on an unauthenticated stdio server would not create authentication. Over stdio there are no HTTP bearer headers: access is normally controlled by the local process boundary, configuration, filesystem permissions, and the spawning client.
Authorization: what can they do?
Once authenticated, what tools can the caller use, and on what data?
function authorize(caller: Caller, tool: string, params: any): boolean {
// Caller-level: can this caller use this tool at all?
if (!caller.tools.includes(tool)) return false;
// Data-level: is this caller authorized for this specific data?
if (params.tenant_id && params.tenant_id !== caller.tenant_id) return false;
return true;
}
Don’t let the LLM make authorization decisions. The LLM might be tricked. Authorization is the server’s job; the LLM only sees data it’s authorized to see.
For multi-tenant systems: every tool call is scoped to a tenant. The tenant is determined by the auth, not by parameters the LLM provides.
Pattern 6: Idempotency
For write operations, idempotency matters. The LLM might retry; it might call the same tool twice in different contexts. Without idempotency, you get duplicates.
Idempotency keys.
The tool accepts an idempotency_key. The server must claim that key atomically in durable storage and bind it to the authenticated principal, tool name, and a hash of the normalized request. A separate read followed by a write races under concurrency.
async function createInvoice(params: {
amount: number;
customer_id: string;
idempotency_key: string;
}) {
return database.transaction(async (tx) => {
const claim = await tx.claimIdempotencyKey({
principal_id: currentPrincipal.id,
tool: "create_invoice",
key: params.idempotency_key,
request_hash: hashCanonicalRequest(params),
});
if (claim.request_hash_mismatch) throw new Error("Idempotency key reused for different input");
if (claim.completed_response) return claim.completed_response;
const invoice = await tx.createInvoice(params);
await tx.completeIdempotencyClaim(claim.id, invoice);
return invoice;
});
}
For the LLM, hint at this in the tool description:
"For each unique invoice you create, generate a UUID and pass it as idempotency_key. If you need to retry the operation, use the same UUID to avoid duplicate creation."
Pattern 7: Streaming
For long operations, first decide whether the operation should be a durable asynchronous job. MCP progress notifications are useful only when the client supplies a progress token and maintains the connection. The exact handler API changed between SDK majors; copy the example for your pinned release rather than this article.
See the pinned SDK’s server example for progress and task APIs. Never emit a progress notification with an undefined token, and never use a transient notification as the only record of a consequential operation.
Use streaming for:
- Operations for which intermediate progress is meaningful to the client.
- Large outputs (so the LLM can start processing while output is still coming).
- Operations with intermediate results worth showing.
Don’t stream for fast, simple operations — adds complexity without value.
Pattern 8: Caching
Many tool calls hit the same data repeatedly. Caching can dramatically improve performance and reduce backend load.
Local cache. In-process cache (e.g., LRU) for hot data.
Distributed cache. Redis or similar for shared cache across server instances.
Cache invalidation. When data changes, evict relevant entries. (This is the hard part.)
TTLs. Cached entries expire after a defined time. Tune per data type — customer profiles might cache for hours; pricing might cache for minutes.
For caching to help, the same authorized calls must repeat. Include tenant and authorization scope in cache keys, and do not cache sensitive responses across callers.
A pattern:
async function getCustomerCached(id: string) {
const cached = await cache.get(`customer:${id}`);
if (cached) {
metrics.increment("cache.hit");
return cached;
}
metrics.increment("cache.miss");
const customer = await db.getCustomer(id);
await cache.set(`customer:${id}`, customer, { ttl: 300 });
return customer;
}

Pattern 9: Rate limiting
LLM agents can be surprisingly aggressive — looping, retrying, fanning out. A misbehaving agent can DoS your backend.
Rate limiting by authenticated principal, client, tool, and resource cost is essential for a remote service. In a v1 tool handler, read the identity from validated extra.authInfo; do not use a nonexistent context.caller:
const limiter = new RateLimiter({
windowMs: 60_000,
max: 100 // 100 calls/minute per caller
});
server.registerTool("expensive_report", {
description: "Build an authorized report.",
inputSchema: { report_id: z.string() },
}, async ({ report_id }, extra) => {
if (!extra.authInfo) return toolError("Authentication required");
const subject = String(extra.authInfo.extra?.sub ?? extra.authInfo.clientId);
if (await limiter.exceeded({ subject, tool: "expensive_report" })) {
return toolError("Rate limit exceeded; retry later");
}
return buildAuthorizedReport(subject, report_id);
});
Beyond global rate limits, per-tool limits matter — some tools are expensive and should be limited tightly.
For consequential operations (creating records, sending messages), use stricter limits or require explicit confirmation flows.
Pattern 10: Resources
MCP has “resources” — read-only data sources the LLM can browse and reference. Different from tools (which are called actively).
server.setRequestHandler(ListResourcesRequestSchema, async () => ({
resources: [
{
uri: "doc://my-server/handbook",
name: "Employee Handbook",
mimeType: "text/markdown",
description: "Company employee handbook"
},
// ...
]
}));
server.setRequestHandler(ReadResourceRequestSchema, async (request) => {
const content = await loadResource(request.params.uri);
return { contents: [{ uri: request.params.uri, mimeType: "text/markdown", text: content }] };
});
Resources are useful for:
- Reference documents the LLM might want to browse.
- Configuration or context data.
- Lookup tables or schemas the LLM might need.
Resources are read; tools are actions. Use the right concept for each.
Pattern 11: Observability
For more on call-level and trace-level patterns for LLM apps generally, see Observability for LLM apps. For your MCP server, instrument:
- Every tool call: timestamp, authenticated subject or pseudonymous identifier, tool, redacted argument summary, result classification, latency, and status.
- Per-tool metrics: call volume, p50/p95 latency, error rate.
- Per-caller metrics: who’s calling, how often.
- Trace context: propagate trace IDs from the caller through to backend calls.
Application-level pseudocode (not an SDK handler context):
logger.info("tool_call", {
tool: request.params.name,
caller_id: authenticatedPrincipal.id,
trace_id: currentTraceId,
params: redactPII(request.params.arguments),
duration_ms: duration,
status: "success"
});
Pipe to your observability platform.
Pattern 12: Versioning
Your MCP server will evolve. Tools will change. New tools added. Old tools deprecated.
Server versioning. The McpServer constructor takes a version. Bump it on changes. Clients can detect.
Tool versioning. When a tool’s signature changes incompatibly, version it: search_customers_v2. Keep the old version available for a deprecation period.
Schema evolution. Add optional fields safely. Removing fields or changing types is breaking.
Deprecation. When deprecating a tool, mark it in its description: “DEPRECATED: use search_customers_v2 instead.”
For production MCP servers used by multiple clients, versioning is essential. Internal-only servers can be more flexible.
Pattern 13: Testing
How do you test an MCP server?
Unit tests. Each tool’s logic, with mocked dependencies. Standard TypeScript testing.
Schema tests. Schemas validate as expected. Edge cases (missing fields, wrong types) handled correctly.
Integration tests. Spin up the server, send actual MCP requests, verify responses. The @modelcontextprotocol/sdk includes test utilities.
End-to-end with a real LLM. The hardest but most valuable. Have an LLM use your MCP server to perform realistic tasks. Verify the LLM uses the tools correctly. Find tool description issues.
An end-to-end test setup (pseudocode; the exact client wiring depends on which LLM client you use — Anthropic’s TypeScript SDK, OpenAI’s, or a framework that supports MCP):
// Start your MCP server as a child process or in-memory transport.
const server = await startTestServer();
// Drive an LLM with the MCP tools attached. The exact API depends on the client.
const result = await runAgent({
mcpServer: server,
systemPrompt: "You are a customer service agent...",
userMessage: "Find the customer Alice and check her open tickets",
});
// Inspect the tool calls captured by the server during the run.
expect(server.callLog.map((c) => c.name)).toEqual([
"search_customers",
"list_tickets",
]);
End-to-end tests catch tool description issues that unit tests can’t.
Pattern 14: Deployment
Where does your MCP server live?
Stdio (local). The server runs as a process; the client invokes it. Best for desktop apps (Claude Desktop, Cursor) and local tools.
Streamable HTTP (remote). The server is a network service. Best for hosted services, shared infrastructure, multi-client access.
For production servers:
- Streamable HTTP is the standardized remote transport candidate; verify client support, session design, authorization, origin/host protection, proxies, timeouts, and scaling behavior.
- Deploy like any web service: containers, load balancing, auto-scaling.
- TLS required.
- Health checks for the deployment platform.
- Graceful shutdown for in-flight requests.
Pattern 15: Security considerations
MCP servers expose capabilities to LLMs. LLMs can be manipulated. Security implications:
Prompt injection through tool inputs. A user’s request might contain text that tries to trick the LLM into calling tools harmfully. Defenses:
- Tool descriptions clear about expected use.
- Authorization on the server side (independent of LLM-decided params).
- Confirmations for consequential actions.
Data exfiltration. Tools that return data can be abused — the LLM might be tricked into returning sensitive data inappropriately. Defenses:
- Authorization checks.
- Logging of what data is accessed by whom.
- Pattern detection for unusual access patterns.
Resource exhaustion. Tools that consume backend resources can be abused. Defenses:
- Rate limiting.
- Resource limits per tool call.
- Circuit breakers when backend is degraded.
Injection into tool outputs. A tool’s output might contain text that, when read by the LLM, manipulates it. Defenses:
- Sanitize outputs where possible.
- Be wary of tools that return user-generated content.
These are real attack surfaces. Treat MCP servers like any production API: defense in depth.
Assembly sketch: not a complete HTTP server
The following excerpt shows how tenant scoping belongs inside each tool. It is intentionally incomplete: db, cache, logger, authenticate, the HTTP route, the bearer-auth middleware, lifecycle management, and tests are application code. Do not paste it and call the result deployed.
import { McpServer } from "@modelcontextprotocol/sdk/server/mcp.js";
import { z } from "zod";
import { db, cache, logger, authenticate } from "./infra.js";
const server = new McpServer({
name: "crm-server",
version: "1.0.0",
});
// === Tool: search_customers ===
server.registerTool(
"search_customers",
{
description: "Search customers by name, email, or company.",
inputSchema: {
query: z.string().describe("Name, email, or company"),
limit: z.number().int().min(1).max(50).default(10),
},
},
async ({ query, limit }, extra) => {
const auth = await authenticate(extra);
const cacheKey = `search:${auth.tenant_id}:${query}:${limit}`;
const cached = await cache.get(cacheKey);
if (cached) return cached;
const customers = await db.searchCustomers({
tenant_id: auth.tenant_id,
query,
limit,
});
const result = {
content: [{
type: "text" as const,
text: JSON.stringify({
customers,
total_matches: customers.length,
truncated: customers.length === limit,
message:
customers.length === limit
? `Showing first ${limit}; there may be more matches.`
: `Found ${customers.length} customer(s).`,
}),
}],
};
await cache.set(cacheKey, result, { ttl: 60 });
logger.info("search_customers", { tenant: auth.tenant_id, query, results: customers.length });
return result;
}
);
// === Tool: get_customer ===
server.registerTool(
"get_customer",
{
description: "Fetch a single customer by id.",
inputSchema: { customer_id: z.string() },
},
async ({ customer_id }, extra) => {
const auth = await authenticate(extra);
const customer = await db.getCustomer(auth.tenant_id, customer_id);
if (!customer) {
return {
isError: true,
content: [{
type: "text" as const,
text: `Customer ${customer_id} not found. Use search_customers to find by name or email.`,
}],
};
}
return { content: [{ type: "text" as const, text: JSON.stringify({ customer }) }] };
}
);
// === Tool: update_customer_email (with idempotency) ===
server.registerTool(
"update_customer_email",
{
description: "Update a customer's email; pass the same idempotency_key on retry.",
inputSchema: {
customer_id: z.string(),
new_email: z.string().email(),
idempotency_key: z
.string()
.describe("UUID for this update; pass the same value on retry to prevent duplicates"),
},
},
async (params, extra) => {
const auth = await authenticate(extra);
// ... idempotency check, validation, update
return { content: [{ type: "text" as const, text: "ok" }] };
}
);
// ... more tools ...
// Deliberately omitted: authenticated Streamable HTTP route and transport lifecycle.
// Start from the official example for the exact pinned SDK version.
This is a design sketch, not a runnable endpoint. A production implementation still needs an official transport example for the pinned SDK, host-header/DNS-rebinding protection, OAuth resource metadata where applicable, TLS at the deployment boundary, authorization tests, idempotency storage, telemetry redaction, limits, graceful shutdown, and a tested client matrix.
What separates production servers from demos
MCP is a focused protocol; building a production-grade server remains API and security engineering. A standards-compatible service can reduce per-client integration work, but compatibility evidence applies only to the clients, versions, transports, auth path, and tools in the tested matrix.
The patterns that matter: focused tool design, prompt-aware schemas, structured error semantics, robust auth, idempotency, observability, security. Skipping any of these creates an MCP server that fails in production.
Build them in. Test against real LLMs. Iterate on tool descriptions. The result is a service agents can call reliably — without a custom integration for each model.



