Designing prompts for production: system, developer, and user layers
Advanced12 min readPrompt Engineering

Designing prompts for production: system, developer, and user layers

A governance pattern for separating trusted instructions, runtime data, and user input, then versioning, evaluating, deploying, and observing prompts according to risk.

What you should be able to do

Treat system, developer, and user layers as an editorial architecture, then map them to the provider API. Keep instruction authority distinct from runtime data placement, and verify intended behavior with evals, telemetry, staged rollout, and rollback.

Saved only in this browser.
In this article

A prototype may begin with an inline string. Production pressure appears when the prompt needs ownership, review, rollback, data handling, multiple features or locales, or measurable behavior; that can happen before launch and has no universal timeline.

You’ll want to change one part of the instructions but not others. You’ll want different behavior for different customer tiers. You’ll want to A/B test versions. You’ll want to roll back when something breaks. You’ll want to know when the prompt last changed, and why.

A production prompt design makes those choices explicit. The prompt text is one artifact inside a governed release and runtime system.

This article uses a three-layer editorial architecture to explain ownership and change cadence, then shows how to map it to provider APIs. It is a reference design, not a universal wire format. For the security boundary, use the OWASP prompt-injection guidance: separating instructions from data helps review and evaluation, but no prompt template creates an authorisation or isolation boundary.

Production prompt work has two separate artifacts: reusable templates and per-request data. Version templates like code. Treat rendered prompts and model responses as sensitive logs whenever they contain user, customer, or internal data.

Three editorial layers, mapped to an API

This design separates three editorial concerns:

System layer. Relatively stable behaviour, identity, and constraints. Owned by the team accountable for cross-feature AI behaviour.

Developer layer. Per-feature instructions, tool-use policy, and output requirements. Owned by the feature team.

User and runtime layer. The user’s request plus dynamic context such as authorised customer data, conversation history, and retrieved knowledge. Constructed for a request or conversation turn.

Mixing ownership, stable policy, feature instructions, and runtime data can make a change harder to review, evaluate, cache, and roll back. Separate them where that improves control; do not force three API fields when the provider or application has a different representation.

Instruction authority and data placement are related but different. The trust hierarchy answers which instruction wins when messages conflict. Data placement answers where the application carries dynamic or untrusted content. Putting retrieved text in a user or context field does not make it authorised, accurate, or safe. Enforce identity, tenant access, data minimisation, tool permissions, and output validation outside the model.

Separating them is foundational:

┌─────────────────────────────────────┐
│ System layer (relatively stable)     │  Identity, behavior, policy intent
├─────────────────────────────────────┤
│ Developer prompt (per-feature)      │  Feature instructions, tools, format
├─────────────────────────────────────┤
│ User prompt (per-call)              │  User query, context, conversation
└─────────────────────────────────────┘

Provider APIs express instruction and content boundaries differently:

  • OpenAI: In the Responses API, use the top-level instructions parameter or a developer message for application instructions and a user message for user input. Do not assume instructions from a previous response carry over when you manage a multi-turn flow.
  • Anthropic: map the design to the Claude Messages API, its system-instruction mechanism, message roles, and tool definitions; supported role placement can vary by model and platform.
  • Gemini: map it to system_instruction and request content, plus the separate tool configuration used by the selected API.

Use a provider-specific adapter and integration tests. Do not copy role names across APIs and assume equivalent precedence, persistence, or tool behaviour.

Layer 1: The system prompt

Within this editorial pattern, the system layer defines cross-feature role and behaviour. Aim to change it less frequently than feature instructions, but version and evaluate it whenever it does change.

A good system prompt covers:

Identity. Stated role. “You are an AI assistant for [Company], specializing in [domain].”

Voice and style. How it should sound. Specific traits, not vague descriptors.

Required behavioural constraints. What the model should refuse, escalate, disclose, or format. Enforce security, permissions, and irreversible-action controls in application code and downstream systems rather than relying on this text alone.

Behavioral patterns. How it handles common situations. Refusals, escalations, uncertainty.

Safety and compliance. Required disclosures, regulatory rules, content policies.

What it should NOT contain:

  • Feature-specific instructions (“for sales emails, do X”).
  • Dynamic context (“the user’s order history is…”).
  • Tool descriptions (those go elsewhere).
  • Things that change frequently.

A system prompt should be only as long as the evaluated behavior requires. A short prompt can be sufficient and a long one can still omit critical rules; measure instruction conflicts, task quality, latency, and token cost rather than targeting a word range.

A reference template:

You are [name], an AI assistant for [company / context].

## Your role
[2-3 sentences on what you do]

## Voice and style
- [Specific trait 1]
- [Specific trait 2]
- [Specific trait 3]
- Do not [anti-pattern 1]
- Do not [anti-pattern 2]

## Hard constraints
- Never [hard rule 1]
- Never [hard rule 2]
- Always [hard rule 3]

## How to handle uncertainty
- If you don't know something factual: say so explicitly.
- If a user asks for something outside scope: offer what you can help with.
- If a request might cause harm: refuse and explain why.

## Format expectations
- Plain text by default
- Use markdown when displaying code or structured data
- Be concise; do not pad responses with filler

This can be the stable policy-bearing component of the prompt design. Confirm with evaluation and telemetry that its intended behaviour survives feature instructions, long contexts, tool results, and adversarial inputs.

Layer 2: The developer prompt

The developer prompt is feature-specific. Different features have different developer prompts.

A summarization feature’s developer prompt:

Task: produce a summary of the document below.

Requirements:
- 3-5 bullet points
- Each bullet is one complete sentence
- Focus on facts and concrete claims, not impressions
- If the document contains numbers, include the most important ones
- Do not include marketing language or speculation
- If the document is ambiguous about something important, note it

Format: plain markdown bullets, no preamble.

A code review feature’s developer prompt:

Task: review the code diff below.

Output a JSON object with:
- summary: 1-2 sentence overview of the change
- concerns: array of specific issues (each: file, line, severity, description)
- suggestions: array of improvements (each: file, line, suggestion)
- approved: boolean (true if no blocking concerns)

Severity levels:
- "blocker": must be fixed before merge
- "warning": should be addressed but not blocking
- "nit": stylistic, optional

Focus on:
- Logic errors
- Security issues
- Performance issues
- Missing test coverage
- Unclear naming or structure

Skip:
- Formatting (handled by formatter)
- Subjective style preferences

Each feature has its own developer prompt. They’re stored separately, versioned separately, evaluated separately.

Layer 3: The user prompt

The user layer is dynamic. It typically includes:

The user’s actual request. “Summarise this document for me.”

Context the system retrieved. Documents from RAG, customer history, conversation history.

Per-call variables. User name, timezone, language preference, account tier.

This layer is constructed programmatically at call time. The structure usually looks like:

{conversation_history_summary}

{retrieved_context}

User's request: {user_query}

Additional context:
- User name: {name}
- User timezone: {timezone}
- User tier: {tier}

The exact structure depends on the feature and provider. Keep volatile, user-specific, and retrieved data out of reusable instruction templates unless an API contract requires another placement. Wherever the data travels, preserve provenance and apply authorisation, minimisation, delimiting, and validation.

Templating discipline

Production prompt assembly often benefits from templates. Inline concatenation becomes harder to review and test as branches, data fields, and features accumulate.

A simple template system:

from string import Template

SUMMARIZE_TEMPLATE = Template("""
$conversation_summary

Document to summarize:
$document

User's specific instructions: $user_instructions
""")

prompt = SUMMARIZE_TEMPLATE.substitute(
    conversation_summary=summarize_conversation(history),
    document=document_text,
    user_instructions=user_query,
)

More sophisticated: a templating library (Jinja2, Handlebars) with conditionals and partials.

{% if user_tier == "enterprise" %}
You have access to advanced analysis features.
{% endif %}

{% if retrieved_context %}
Relevant context from your knowledge base:
{{ retrieved_context }}
{% endif %}

User's request: {{ user_query }}

Templating keeps prompt structure consistent, enables conditional logic, and makes it easier to delimit or escape user input. It does not, by itself, stop prompt injection — treat untrusted text as data, never as instructions.

Choose a governed source of truth

Treat production prompts as versioned release artifacts. The source of truth can be a repository, an owned registry or service, or a managed product with an editing UI. Choose by governance and operating needs rather than by the idea that one storage pattern fits every team.

For code-managed prompts, a prompts/ directory with one file per prompt is a clear starting pattern:

prompts/
  system/
    main.txt
    customer-support.txt
    code-assistant.txt
  features/
    summarize.txt
    classify-ticket.txt
    generate-email.txt
  templates/
    base.j2

Each file can have its own commit history and PR review, while production deployments reference a known code version.

Why this matters:

  • Diff visibility. When a prompt changes, the diff is in the PR. Reviewers can see exactly what changed.
  • Rollback. When a change breaks something, you can revert.
  • History. “When did we change the refund policy in the prompt?” “Why is this paragraph here?” — answerable via git blame.
  • Tooling. Linters, validators, eval suites all integrate with file-based prompts.

Compare the main options explicitly:

Source of truthGood fitControls to require
RepositoryEngineering-owned prompts released with application codeBranch protection, code owners, sanitised fixtures, environment promotion, deployment identity, rollback to a known commit
Owned registry or serviceRuntime selection, independent prompt releases, multiple products or localesRole-based access control, immutable versions, approval records, environment separation, authenticated clients, encryption, audit events, export and tested rollback
Managed service or editing UINon-engineer collaboration or experiment operationsRole-based access control, least-privilege roles, review workflow, version provenance, production access separation, data-processing review, export and rollback

Inline strings are workable for a small prototype but are harder to discover and release independently. An editing UI can be appropriate when its permissions, review, provenance, deployment, and rollback controls meet the workflow’s risk. Pasted prompts from chat windows should enter the same review, sanitisation, and evaluation path as any other candidate.

Do not commit production conversations, customer records, support tickets, internal documents, or rendered prompts containing sensitive variables. Source control is for reusable templates, fixtures, and sanitized eval examples. Real traces belong in an observability store with retention, access control, and redaction.

Runtime registries and services

When prompt releases need a different cadence from application deployments, or authorised editors need a controlled UI, a repository alone may not provide the right workflow. A registry or service can select an approved version at runtime.

Pattern: a database or service that stores prompt versions with metadata.

prompt = prompt_service.get(
    name="summarize",
    version="v3",
    locale="en",
    user_tier="enterprise",
)

The service should maintain:

  • Current and historical versions of each prompt.
  • Metadata: when added, by whom, why.
  • Eval scores attached to each version.
  • Rollback capability.

The storage engine is only one part of the design. Compare role-based access, audit history, authenticated runtime access, environment promotion, deployment consistency, rollback, export, encryption, data handling, availability, and operating cost. A small database can be sufficient only when the surrounding controls meet the requirement.

Define who may draft, review, approve, release, roll back, and read rendered prompts. Editing access does not imply production-release access. Sensitive workflows may require separate roles, redacted previews, dual approval, or an interface that never exposes live customer data.

Eval-gated changes

Prompt changes capable of affecting user outcomes, tool use, data handling, policy compliance, or downstream decisions should go through review and risk-proportionate evals before deployment. Low-risk copy changes may need a small regression set; a prompt that can influence a payment, account, or regulated decision needs stronger offline cases, adversarial tests, approval, and staged release. Define an emergency path with scoped rollout, monitoring, approval, and rollback rather than silently bypassing the gate.

The flow:

  1. Engineer or non-engineer drafts a prompt change.
  2. The change runs against the eval suite.
  3. Eval results are reviewed alongside the change.
  4. If evals pass (no regressions, ideally improvements), the change can be approved.
  5. Approved changes deploy.
  6. Post-deploy monitoring catches anything evals missed.

For material prompts, keep a representative eval set and run the stable, automatable checks in CI when the signal is reliable enough to gate a change. Use expert or human review where the acceptance criterion cannot be reduced to an automated score. Record what was tested, the threshold, the reviewer, and the residual uncertainty.

This gate does not prove correctness. It makes the release decision inspectable and gives the team a baseline for detecting regressions after deployment.

A revised instruction card waits behind a varied test tray before release
AI-generated illustration of an eval gate that keeps a prompt change out of production until it passes representative tasks.

A practical release checklist

Before a prompt version goes live, require a short checklist:

CheckRequirement
OwnershipPrompt has a named owner and reviewer.
Instruction layersSystem, developer, and user/context data are separated.
SchemaStructured outputs have a schema and failure path.
Injection handlingUntrusted content is delimited, kept outside trusted instruction fields where the API permits, and covered by hostile-instruction tests.
EvalsCandidate prompt passes the regression set and safety cases.
LogsTemplate version, model, latency, cost, and redacted inputs/outputs are observable.
RollbackPrevious known-good version can be restored without code surgery.

The companion checklist linked from this article turns these checks into a repeatable release review.

A/B testing in production

For prompts suitable for online experimentation, a staged comparison against the current version can add real-world signal beyond offline evals. Do not use live traffic as the first safety test, and do not expose people to a materially riskier treatment merely to gather data.

Illustrative pattern, not a default split:

  • 95% of traffic uses production prompt v3.
  • 5% gets new candidate v4.
  • Keep model access, tool permissions, and action limits no broader than the approved production boundary.
  • Measure: task success, safety and policy errors, user feedback, downstream metrics, and reviewed eval scores on eligible traffic.
  • Define the minimum sample, stop conditions, owner, and one-step rollback before starting.
  • After sufficient evidence, decide whether to expand, revise, or stop the candidate.

Delivery mechanisms include owned feature flags, deployment configuration, an approved prompt registry, or custom routing.

Caveats:

  • A/B testing only catches signals you measure. If you don’t have user feedback or downstream conversion metrics, A/B testing tells you little.
  • Statistical significance requires volume. For low-volume features, A/B is hard.
  • Concurrent experiments can interact and confound attribution; control overlaps deliberately.
  • Consider notice, consent, exclusion, and review obligations for the product, population, and jurisdiction. High-impact or irreversible actions usually need stronger approval and reversibility than a traffic split can provide.

Observability of prompts

Define privacy-safe telemetry for each production workflow. For calls that affect material outcomes, capture enough of the following to identify the deployed behaviour and reconstruct failures:

  • Which prompt template was used (name, version).
  • Which variables were substituted, using allowlisted names and redacted values where needed.
  • The final rendered prompt only when policy permits it; otherwise store a redacted, sampled, or hashed representation.
  • The model’s response, redacted or sampled for sensitive workflows.
  • Latency, tokens, cost.
  • Downstream signals (user feedback, success metrics).

The objective is to answer which version ran, what authorised evidence it received, what validated output it produced, which tools or policies were involved, and what happened downstream. Do not log hidden reasoning as a substitute for those facts.

Storage: a database table or observability tool. Costs are real if you log everything (call volume × token count × storage). Privacy risk is also real if you log everything. Decide per workflow which fields are safe to store, redact secrets and personal data by default, and keep retention short unless there is a compliance reason to keep traces longer. Some teams sample.

Set a review cadence and sampling method from volume, risk, change rate, incidents, and legal constraints. Review redacted or otherwise approved traces, include known edge cases, and feed confirmed failures back into evals. A weekly sample may suit one workflow and be inappropriate or insufficient for another.

Prompt anti-patterns

A few patterns to avoid:

Anti-pattern 1: An unowned multi-purpose prompt. A large system prompt that combines unrelated features, policies, examples, and runtime assumptions can be hard to review, evaluate, and roll back. Length alone is not the defect; unmeasured complexity is.

Fix: separate components by ownership and change boundary where useful, remove duplicated or obsolete instructions, and compare the revised design on representative evals.

Anti-pattern 2: Inline string concatenation.

prompt = "You are helpful. " + (
    "The user is a paid customer. " if user.tier == "paid" else ""
) + f"Their name is {user.name}. " + ...

Fragile and hard to read. String concatenation also makes instruction/data boundaries difficult to inspect, but moving the same values into a template does not neutralize malicious instructions.

Fix: templating system.

Anti-pattern 3: Same prompt for too many use cases.

A single “general assistant” prompt used for email drafting, code review, customer support, and research can hide task-specific acceptance criteria and ownership.

Fix: feature-specific developer prompts on top of a shared system prompt.

Anti-pattern 4: Hard-coded prompts.

response = openai.chat.completions.create(
    messages=[
        {"role": "system", "content": "You are a helpful assistant..."},
        {"role": "user", "content": query}
    ]
)

The prompt is buried in the code. Can’t be edited without a deploy. Can’t be A/B tested. Can’t be versioned independently.

Fix: extract to a prompt file or service.

Anti-pattern 5: No eval coverage.

A feature ships with a prompt that’s never been tested systematically. Quality is “vibes.” Drift is undetectable.

Fix: add risk-proportionate eval coverage for the behaviours that matter, including failure and escalation cases.

Anti-pattern 6: Mixing data into the system prompt.

You are an assistant for John, a premium customer who joined in 2023, lives in Tallinn, and has 47 open tickets.

Now the reusable instruction prefix changes on each call, which can reduce cache reuse and obscure the distinction between policy and customer data.

Fix: carry dynamic data in the provider’s runtime input or context mechanism, with provenance, authorisation, and minimisation. Do not infer trust from its message role.

Anti-pattern 7: Instructions buried in the middle.

Help the user with their request. Be polite. Format output as JSON. Don't use markdown. The user is asking about pricing, so be careful about quoting numbers. Output should be 1-2 sentences. Now help them.

Critical requirements are harder for a reviewer to discover and can conflict with nearby prose.

Fix: group critical instructions in a clearly labelled block, state each rule once, and test whether the selected model follows them under representative long and adversarial contexts.

Specific patterns for common features

A few feature-specific patterns:

Classification

Task: classify the following text into one of these categories:
- billing: payment, refund, subscription
- technical: bug, error, integration issue
- account: login, password, profile changes
- feature_request: new functionality requests
- complaint: general dissatisfaction without specific actionable issue

Output a JSON object: {"category": "<one of above>", "confidence": "<high|medium|low>", "reasoning": "<1 sentence>"}

Text to classify:
{text}

Patterns: enumerated categories with definitions, structured output, confidence and reasoning fields.

Extraction

Task: extract structured data from the document below.

Schema:
- vendor_name: company that issued the invoice
- invoice_number: as printed on the document
- date: ISO 8601 format
- line_items: array of {description, quantity, unit_price, total}
- subtotal, tax, total: numbers

Rules:
- If a field is not present, use null
- Numbers should be numeric, not strings
- For ambiguous cases, set "needs_review": true and explain

Document:
{document}

Patterns: explicit schema, type expectations, handling of missing data, escalation for ambiguity.

Generation with style

Task: write a {format} on the topic of {topic}, targeting {audience}.

Style:
- {Specific style trait 1}
- {Specific style trait 2}
- Avoid: {anti-pattern 1}, {anti-pattern 2}

Constraints:
- Length: {N} words
- Include: {required elements}
- Exclude: {forbidden elements}

Voice reference:
[Provide a sample of the desired voice]

Output: the {format} only, with no preamble or post-script.

Patterns: specific style traits (not generic), explicit constraints, voice anchored with reference sample.

Agent loop

You have access to the following tools:
{tool_descriptions}

For each turn:
1. Think about what you need to do.
2. Decide if you need a tool. If yes, call it.
3. After observing the result, decide if you need more tools or can answer.
4. When you have enough information, produce the final answer.

Constraints:
- Maximum 5 tool calls per request.
- If after 5 calls you can't complete, explain what's missing.
- Never invent tool names or arguments.
- Verify tool results before acting on them.

User request:
{user_query}

Patterns: a staged tool-use procedure, a tool budget, explicit validation, and bounded failure handling.

The team aspect

Prompts in production usually involve multiple people:

  • Engineers wire the prompts into the system, maintain templates, manage deployments.
  • Product defines what the prompts should accomplish.
  • Content/marketing owns voice and style guidance.
  • Domain experts know what’s right for specific use cases (legal language, medical terms, etc.).

A useful pattern: a “prompt review” process similar to code review, with the right reviewers for each domain. Voice changes get content’s review. Logic changes get engineering’s review. Domain-specific content gets the expert’s review.

For sensitive use cases (legal, medical, financial), prompts may need formal review and sign-off. Build the process accordingly.

A stage-gated prompt maturity plan

For teams moving from “prompts are strings in code” to “prompts are managed infrastructure”:

Stage 1: Foundation.

  • Inventory the prompts and prompt-building code that materially affect behaviour.
  • Define the instruction authority model, runtime-data boundary, owners, and provider mapping.
  • Choose a governed source of truth and a templating approach appropriate to the release workflow.
  • Set up privacy-safe telemetry that identifies the deployed version and outcome without retaining unnecessary sensitive content.

Stage 2: Evaluation.

  • Build eval suites for the highest-risk and highest-volume prompts first.
  • Run representative evals on material prompt changes, with domain review where needed.
  • Put stable automated checks in CI and keep non-automatable acceptance decisions visible in the review record.

Stage 3: Operations.

  • Implement prompt versioning in the chosen repository, registry, or service.
  • Add a staged rollout and stop mechanism; use A/B testing only where the workflow is eligible.
  • Build monitoring for the quality, safety, policy, latency, cost, and downstream signals the workflow requires.
  • Establish review process for prompt changes.

Do not promise this outcome on a calendar. Exit the sequence only when tests show prompt discovery is complete, critical changes are versioned and gated, rollback works, telemetry identifies the deployed version, and owners can rehearse an incident response.

Prompts as infrastructure

Production prompts may be stored as strings, but they operate as versioned system components with surrounding controls for assembly, evaluation, release, access, and observability.

An instruction hierarchy defines authority; it does not secure runtime data or tool actions. Templating can reduce assembly mistakes but does not prevent prompt injection. A governed source of truth provides version history. Risk-proportionate evals inform release decisions, and telemetry tests whether intended behaviour persists outside the eval set.

The required rigor depends on risk and scope, but any omitted control needs a recorded rationale. Publication claims should show the versioning, evaluation, rollout, rollback, and telemetry evidence actually implemented.

Intended controllability is a hypothesis until evals and production telemetry show that the selected model, prompt version, data path, tools, and policies behave within the acceptance thresholds. Preserve that evidence with the release, investigate failures by version, and keep rollback usable.

Start with the smallest architecture that makes ownership, authority, data handling, evaluation, deployment, rollback, and evidence explicit.

Read next

Continue through the same learning path with the next practical articles.