When a description leaves too much room for interpretation, show the model examples of the output you want. A small set of representative input-output pairs can make tone, format, or classification boundaries more concrete than adjectives alone.
This technique is called few-shot prompting. The name comes from machine-learning research: the original GPT-3 paper studied how models perform when a task and examples are provided in context, without task-specific parameter updates (Brown et al., 2020). The paper also found tasks where few-shot performance struggled, so examples are a technique to evaluate, not a universal upgrade.
This article walks through what few-shot prompting is, when to use it, and how to test it with worked examples.
Why examples work better than descriptions
Try this thought experiment. Suppose you want someone to match your company’s tone of voice for marketing copy. You could say “friendly but professional, warm but not casual, confident but never arrogant, plain English but not dumbed down.” After those four phrases, they would probably nod and still produce something different from what you actually wanted.
But if you instead showed three short paragraphs that already hit the tone, they would have a more concrete target. They could compare word choice, sentence length, and structure instead of interpreting the adjectives alone.
Models can use examples this way too. Adjectives like “friendly,” “professional,” and “confident” leave many valid interpretations. Concrete examples expose choices the instruction may not name. OpenAI’s prompt engineering guide describes few-shot learning as steering a model with a handful of input-output examples and recommends showing a diverse range of possible inputs with their desired outputs. Whether that improves your result depends on the model, task, examples, and evaluation criteria.
When to reach for few-shot
Some situations where examples are worth testing:
Matching a specific tone. “Write in our company voice.” If a short description leaves room for interpretation, show representative accepted examples too.
Producing consistent formatting. Anything where the output needs to look the same every time, such as product descriptions, error messages, API responses, weekly reports, or status updates. Show the format as well as describing it.
Niche or unusual outputs. “Write a thread for X the way [a specific person you follow] writes them.” “Generate headlines for an internal tool blog.” “Write code comments the way our team writes them.” These have specific conventions that are hard to articulate.
Replicating a style for translation, summarisation, or rewriting. “Rewrite this in the same style as these three examples I’ll show you.”
Anything you keep correcting in the same way. If you find yourself editing the model’s output in the same direction every time, such as making things shorter, swapping a word, or tightening structure, give it examples of the fixed version.
The basic structure
A practical few-shot prompt has three parts: a brief instruction, a representative example set, and the new task.
Generate a one-line product description for a B2B SaaS tool. Match the style of these examples.
Example 1: Product: ProjectHub Description: A shared workspace for project teams who are tired of switching between five tools to do one job.
Example 2: Product: TimeFlow Description: A time-tracking app for people who hate time-tracking apps.
Example 3: Product: ClearStack Description: A reporting tool that turns spreadsheets into decisions.
Now write one for: Product: PromptDesk Description:
The requested pattern is short, opinionated, slightly cheeky, and structured around a specific complaint or audience. Check the result against those criteria; examples can guide the pattern, but they do not guarantee a good description.
Three worked examples
Let’s walk through three situations where examples can make the target clearer.
1. Bug report comments in your team’s style
Suppose your team writes Jira tickets in a specific way: short, focused on user impact, and free of jargon. You want AI to draft tickets that match.
Draft a Jira ticket description following the format below.
Example 1: User logging in via Google sees a brief flash of the wrong language before the UI updates. Happens consistently on Chrome desktop. Not blocking, but feels janky.
Steps:
- Sign out
- Sign back in with Google
- Observe initial page flicker
Expected: language stays consistent Actual: brief flash of (apparent default) English
Example 2: Export to CSV button returns an empty file on weekly reports >1000 rows. Smaller reports export correctly.
Steps:
- Open weekly report with 1000+ rows
- Click Export → CSV
- Open downloaded file
Expected: full data Actual: file is 0 bytes
Now draft one for: Issue: Users on Safari say the mobile menu won’t close after they tap a link. Refresh fixes it. Looks like a focus-trap issue.
Review the draft for the requested structure, tone, and factual fidelity. The examples make the target explicit, but the model can still omit a field or infer a detail that was not supplied.
2. A brand-voice rewrite
Suppose you have a draft email and want it rewritten to match a brand voice. “Be warmer” is open to interpretation; examples can make the intended register more specific.
Rewrite the email below to match the voice of these examples. The voice is: direct, no corporate filler, slightly self-aware, never says “synergy” or “leverage.”
Example 1: “We pushed the release back by a week. The autoscaling change was bigger than we thought. New ship date: Friday the 22nd.”
Example 2: “Quick favour: could you sanity-check this draft? Specifically the second section. I think it’s overreaching but I can’t tell.”
Example 3: “Heads up: I’m going to push back on the deadline in tomorrow’s meeting. The math doesn’t work and I’d rather flag it now than miss it.”
Rewrite this in the same voice:
[paste your draft]
Compare the rewrite with the examples and your source facts before using it. If your ChatGPT plan and workspace allow GPT creation, you can put the examples in a custom GPT’s instructions; OpenAI recommends testing configured GPTs in Preview because instructions do not guarantee identical output on every run.
3. Structured extraction
Suppose you have approved invoices in PDF and want an eligible model to extract them into a clean format. After confirming the tool and data flow are approved for the documents, upload one, provide an example of the desired output, then ask for the rest.
Extract data from invoice PDFs in this exact JSON format.
Example:
Input: [invoice PDF where vendor is “Lufthansa”, date is 2026-04-12, total is 423.50 EUR, line items are flight + bag fee]
Output:
{ "vendor": "Lufthansa", "date": "2026-04-12", "currency": "EUR", "total": 423.50, "line_items": [ {"description": "Flight TLL-LHR", "amount": 387.00}, {"description": "Checked bag", "amount": 36.50} ], "category": "Travel" }Now extract from this invoice: [attach new PDF]
The example specifies the target schema, but you still need schema validation and checks against the invoice. Do not treat a plausible JSON object as evidence that amounts, dates, or line items were extracted correctly.

Common mistakes with few-shot
Three mistakes to watch for:
Examples that contradict each other. If examples vary in style or label similar inputs differently, the model may reproduce the ambiguity. Keep the intended rule consistent while covering meaningful input variation.
Examples that do not cover meaningful variation. There is no universal best count. Start with the smallest set that demonstrates the pattern and important edge cases, then add or remove examples based on representative tests. Every example consumes context and can introduce another pattern for the model to reconcile.
Examples that include errors you do not want copied. A model can reproduce typos, awkward phrasing, or accidental structural choices from its examples. Clean the set before using it and check new outputs rather than assuming only the intended pattern transferred.
A particularly subtle mistake: showing examples of input but not output. “Here are three articles I want you to write headlines for: [three articles].” That is zero-shot, not few-shot. The model has not seen what a good headline looks like. Compare:
Here are three articles with the kind of headlines I’d like. Match the style:
Example 1: [article] → Headline: ”…” Example 2: [article] → Headline: ”…” Example 3: [article] → Headline: ”…”
Now write a headline for: [new article]
The pair (input → output) is the unit of few-shot. Without the output side, the model has no target pattern to follow.
How to build a library of examples
Once you start using few-shot, you will accumulate sets of “examples that worked.” Keep them. Save them in:
- Custom GPTs / Claude Projects: store tested examples in scoped instructions that you can review and update.
- A text snippet manager (TextExpander, Raycast, Espanso): use shortcuts that expand into your examples.
- A notes file organised by use case (“brand voice examples,” “ticket format examples,” “headline examples”) that you can paste from quickly.
Over time, this can become a task-specific library grounded in outputs your team has actually accepted, which is more useful than a generic prompt pack that has not been tested on your work.
Examples beat adjectives
When you can describe the output precisely, describe it. When the style is something you recognise but cannot fully articulate, add representative input-output examples. Then test the prompt on ordinary cases and edge cases before turning it into a reusable workflow.
Examples can clarify what adjectives leave ambiguous. Try the technique on a prompt where style or format matters, and keep it only if the tested outputs improve.



