Three studied prompting/search patterns are chain-of-thought, self-critique, and tree-of-thoughts. Chain-of-thought was evaluated in Wei et al., 2022 and tree-of-thoughts in Yao et al., 2023, on particular models and benchmarks. Those papers do not establish a universal improvement for current models or business tasks. Treat each pattern as a hypothesis to compare against a direct prompt on your own eval set.
Here is what each technique is, when to use which, and the cost-benefit math.
What problem these solve
All three techniques address the same root issue: a language model normally generates an answer autoregressively, one token at a time. Without a separate deliberation or search phase, early tokens quickly constrain what follows and can lock the answer onto a weak conclusion. For simple tasks, this is fine and efficient. For multi-step reasoning, complex analysis, or anything where the answer depends on getting several intermediate things right, that default can produce confidently wrong outputs.
The prompts request intermediate structure or additional passes. More visible text or more calls can increase cost without establishing that the underlying answer is correct.
Chain-of-thought (CoT)
The original pattern asks a model to generate intermediate reasoning before a conclusion. For a real workflow, prefer requesting the verifiable artifacts you need — equations, assumptions, cited evidence, test output, or a concise decision rationale — rather than treating free-form internal narration as evidence.
A worked example. Compare:
Vanilla: A train leaves Tallinn at 9:00 a.m. traveling 80 km/h. Another train leaves Tartu at 9:30 a.m. traveling toward Tallinn at 100 km/h. The distance between the cities is 190 km. At what time do they meet?
versus:
With CoT: Same question. Think step by step. First, calculate the distance the first train covers before the second starts. Then set up the equation for when they meet. Show your work, then give your final answer.
The 2022 paper reported improvements on several arithmetic, commonsense, and symbolic-reasoning benchmarks for sufficiently large models, with material variation by task and model. Do not transfer those historical results to a current model, a different prompt, or production data. Run both prompt variants and score the final answer and checkable intermediate artifacts.
When to use chain-of-thought:
- Multi-step arithmetic, especially with units, dates, or precise rounding. Even strong models slip on these.
- Logic puzzles and similar problems where the answer is the conclusion of a chain.
- Code debugging, where the answer depends on tracing through state.
- Strategy analysis, where the conclusion depends on weighing multiple factors.
When not to bother:
- Simple factual recall. “What is the capital of Estonia?” needs no CoT.
- Generation tasks. Writing, summarising, drafting. CoT just adds tokens without quality gain.
- Reasoning models. o3, Claude Extended Thinking, DeepSeek R1 already do CoT internally — adding “think step by step” to your prompt is at best redundant, at worst counterproductive.
That last point is critical and we’ll return to it.
Self-critique
A two-pass technique. First, ask the model to produce an answer. Then, ask the model to critique its own answer and produce a revised version.
The prompt structure:
Step 1: [Your original question]
Step 2: Review your answer above. Find any mistakes, weaknesses, or places where you made assumptions that might not hold. Be a hard critic of your own work.
Step 3: Based on your critique, produce a revised answer.
A critique pass can catch some defects, but the same model can repeat or rationalise the same error. Treat self-critique as another fallible evaluator; use deterministic checks, retrieved evidence, tests, or independent qualified review where consequence warrants it.
A more sophisticated variant is constitutional / principle-based self-critique. You define a set of principles the answer should satisfy, then ask the model to evaluate against each.
Principles for a good answer to this kind of question:
- It addresses the actual question, not a generalisation of it.
- It quotes specific evidence rather than gesturing at sources.
- It acknowledges uncertainty explicitly where present.
- It is calibrated — confident on strong points, hedged on weak ones.
Produce an answer. Then evaluate it against each principle. Then revise.
This is the technique behind Anthropic’s Constitutional AI work and similar approaches in modern alignment research.
When to use self-critique:
- Writing tasks where you want a second pass without leaving the conversation.
- Analytical work where the model is likely to be overconfident.
- Decision support where you want the model to find the holes in its own argument.
- Code where you want a review pass after the generation pass.
When not to bother:
- Tasks where there is no “correct” answer to revise toward (creative brainstorming, idea generation).
- Tasks where you would prefer to do the critique yourself (anything where your judgement is the value).
- Quick conversational responses where the latency cost outweighs the quality gain.
Tree-of-thoughts (ToT)
The most expensive technique. Instead of producing a single chain of reasoning, the model explicitly considers multiple paths, evaluates each, and selects the most promising.
A worked example structure:
Step 1: Generate three different approaches to this problem.
Step 2: For each approach, work through the first few steps without committing to a final answer.
Step 3: Evaluate which approach is most likely to succeed and why. Be specific about strengths and weaknesses.
Step 4: Commit to the best approach and complete the solution.
Tree-style search is a candidate when a problem has multiple plausible approaches and the workflow can evaluate partial paths. Generating three options does not ensure diversity or avoid a shared false assumption; define evaluation criteria and retain the evidence for the selected path.
Practical example — a hard prompt:
I have a complex SQL query that runs too slowly. Help me optimise it.
Step 1: Generate three different optimisation strategies. Step 2: For each, identify the specific bottleneck it would address and the cost. Step 3: Evaluate which is most likely to give us the biggest gain for the smallest risk. Step 4: Implement the chosen approach.
You get back something noticeably more thoughtful than “here is one rewrite.” Three approaches, comparison, recommendation, implementation.
When to use tree-of-thoughts:
- Problems with multiple credible solutions. Architecture decisions, algorithm choices, strategic choices.
- Optimisation problems. Where the first attempt is rarely the best.
- Creative tasks where exploration is the point. Naming, framing, positioning.
- Anything where you suspect the obvious answer is wrong.
When not to bother:
- Tasks with a single correct approach. Don’t ask for three SQL queries when one works.
- Simple factual questions. Overkill.
- Most reasoning-model tasks — the models do this kind of exploration internally now.
A practical decision tree
When you have a hard problem in front of you, the question is not “should I use CoT, self-critique, or ToT.” It is “what is the shape of the problem?”
- Linear multi-step problem (arithmetic, logic puzzle, strict reasoning) → chain-of-thought.
- Problem where overconfidence is the risk (analysis, recommendation, code that should be reviewed) → self-critique.
- Problem with multiple plausible approaches (optimisation, strategic choice, creative exploration) → tree-of-thoughts.
- Conversational, simple, or generative → none. Skip the overhead.
Do not stack techniques by default. Compare a direct prompt, one scaffold, and any combined variant on the same cases. Keep the simplest version that meets the release criteria.
How reasoning models change the prompting test
Reasoning-capable models expose controls and behaviours that differ by provider and version. Model names, retirement dates, reasoning settings, and visibility of intermediate content change; use the current provider documentation instead of the frozen list in an article.
This changes how to prompt them in three important ways:
1. Start outcome-first. State the objective, important context, constraints, evidence requirements, and completion bar. OpenAI’s current GPT-5.6 prompting guidance recommends describing the destination and validating prompt simplifications on representative tasks.
2. Verify the result, not the amount of reasoning. More reasoning effort or latency is not proof of correctness. Require citations, calculations, schema validation, tests, or other task-appropriate evidence and measure the final result.
3. Compare a direct prompt with necessary structure. Remove redundant instructions one group at a time, but keep real constraints and required output schemas. A shorter prompt is better only when it still passes the same evals.
A worked example. Compare these two prompts to a reasoning model:
Prompt A: “Think step by step about the following question. First, identify the key constraints. Then list the options. Then evaluate each option against the constraints. Then choose. Show your reasoning at each step. Question: should we adopt a four-day work week?”
Prompt B: “Should we adopt a four-day work week? Context: 80-person B2B SaaS, customer support team operates Mon-Fri.”
Prompt B is a useful baseline, but it omits decision criteria, evidence, stakeholder impacts, and a completion bar. Compare it with a concise structured prompt and score both; do not declare a winner from prompt length alone.
Different models and settings can respond differently to scaffolding. Store prompts with model/version and eval results, then retest during upgrades.

Cost-benefit tradeoffs
The techniques can add generated tokens, calls, and latency. There is no portable multiplier without a fixed model, task, prompt, reasoning setting, and stopping condition:
| Technique | What to measure | Candidate use |
|---|---|---|
| Requested intermediate work | final accuracy, artifact correctness, tokens, latency | tasks with checkable intermediate artifacts |
| Self-critique | defects caught, new defects introduced, false confidence, second-pass cost | revision against an explicit rubric |
| Multiple-path search | path diversity, evaluator accuracy, total calls, tail latency | problems with genuinely different candidate approaches |
| Provider reasoning mode | task pass rate by effort setting, tokens, latency, price | only where a higher effort setting improves the release eval |
Use a direct prompt as the baseline. Add a scaffold or reasoning setting only when the measured quality improvement is worth the additional cost and latency for that task.
A practical rule: before you reach for a technique, ask whether the cost of being wrong on this task is meaningful enough to justify the extra cost of the technique. If yes, use the right one. If no, just send the prompt.
Worked example: a real hard task
Suppose you are evaluating two vendor proposals and you want a calibrated comparison.
Without any technique (vanilla prompt):
Compare these two vendor proposals [paste]. Which should we pick?
You get a hedged, both-sides answer. Useful starting point; not enough.
With CoT:
Compare these two vendor proposals. Think step by step:
- List the criteria that matter for our decision.
- Score each vendor on each criterion.
- Identify the criteria where the scores diverge most.
- Then give your recommendation.
You get a much more structured analysis. Each step is visible; you can verify or correct.
With self-critique on top:
[same as above]
After your recommendation, critique your own analysis:
- Which criteria might I have weighted wrong?
- What did I assume that I shouldn’t have?
- What’s the strongest credible case for the other vendor?
Then produce a revised recommendation if needed.
The critique creates an opportunity to find blind spots; verify whether it actually did.
With ToT:
Compare these two vendor proposals.
Step 1: Generate three different decision-making frameworks for this kind of choice (e.g., risk-minimising, value-maximising, capability-aligned). Step 2: Apply each framework. Get three recommendations. Step 3: Where do the frameworks agree? Where do they diverge? Step 4: Given our actual constraints, which framework is most appropriate? Final recommendation.
The prompt requests three angles. Check that they are substantively different and evidence-based rather than rewordings of one assumption.
With a reasoning model:
Compare these two vendor proposals. Which should we pick, and why? Include the things that would change your answer.
This is the direct baseline. The visible answer does not prove which internal process occurred. Compare its correctness, evidence, tokens, and latency with the structured variants.
For hard analytical work, evaluate a current reasoning-capable model and a direct outcome-first prompt before adding scaffolding. Keep any scaffold only when it improves the task’s eval.
A few practical habits
Make the technique visible to yourself. Note which technique you used in your prompt — it helps you build intuition about what works.
Compare outputs. Run the same representative cases with and without the technique after relevant model, prompt, data, or policy changes. Score against a rubric; visual difference alone is not quality.
Don’t stack techniques without thought. Stacking CoT + self-critique + ToT + reasoning model is rarely better than picking the right one. Each layer adds cost; only add layers that genuinely improve the answer for your specific task.
Keep the techniques in your library. Snippets for “with CoT,” “with self-critique,” “with ToT” — applied to whatever the current task is — save real time over re-typing the scaffolding.
Pick the technique that earns its cost
These techniques are candidates, not guarantees: intermediate artifacts for checkable multi-step work, self-critique for a rubric-based second pass, and multiple-path search for problems with genuinely different approaches. Reasoning-capable models change the comparison, so re-evaluate rather than carrying forward historical prompt folklore.
Use the right one for the problem. Skip them when they don’t earn their cost. Distinguish “harder problem” from “different problem” — the technique choice follows from that.



