Providers expose reasoning in different ways. OpenAI uses reasoning effort and modes on supported models, Anthropic documents adaptive and extended thinking, and Google documents thinking levels or budgets for supported Gemini models. Product labels and parameters change, so choose by current provider documentation and measured task performance rather than a fixed list of model names.
Some reasoning models benefit from a different prompting style from non-reasoning configurations. For OpenAI reasoning models, the current guidance is to start with straightforward prompts and avoid asking for a chain of thought. Treat broader claims about scaffolding, roles, and self-critique as hypotheses to test on your model and workload, not rules that transfer automatically between providers.
This covers what reasoning models are, when to use them, how to prompt them well, and the pitfalls that catch even experienced AI users.
What reasoning controls change
Providers use labels such as reasoning, thinking, effort, and budget for controls that can change token use, latency, and output quality. Depending on the provider and configuration, reasoning content may be hidden, summarised, omitted, or returned in separate blocks. Do not infer a provider’s hidden process from a label or from the style of the final answer.
Reasoning can improve results on some hard, multi-step tasks, but the gain depends on the model, effort setting, prompt, and evaluation. It can also add latency and billed reasoning tokens. The choice between a lower-latency configuration and more reasoning is therefore an architectural decision to benchmark, not a universal upgrade.
Examples of reasoning product families in 2026 include:
- OpenAI reasoning modes. OpenAI has shipped o-series models and GPT reasoning modes; availability and retirement dates vary by product and plan, so check the current model guidance before standardising.
- Claude models with thinking. Anthropic’s current documentation says thinking is available on current Claude models, but the supported configuration varies: some use adaptive thinking, while manual
budget_tokensis unsupported or deprecated on others. Follow the per-model table rather than applying one generic parameter recipe. - DeepSeek R1. The R1 paper describes the model family and training approach. Check DeepSeek’s current official documentation for supported models and access methods.
- Gemini thinking modes. Supported Gemini API models expose provider-specific thinking controls documented in the Gemini thinking guide.
- Other provider offerings. Treat product names, entitlements, and controls as time-sensitive. Verify them in the provider’s current official documentation before adopting or recommending them.
These products differ in model behavior, supported controls, cost, latency, and how much reasoning output is visible. Do not assume that a parameter or prompting rule transfers unchanged between providers.
When to test more reasoning
A reasoning model or higher-effort setting can be a relevant comparison. Use these task signals and boundaries when defining the experiment:
- The problem has multiple steps that build on each other. Math, logic, multi-step planning, code that requires tracing state.
- A lower-reasoning configuration keeps failing the same evaluated cases. A reasoning model or higher effort is then a relevant experiment. Compare both configurations on the same cases.
- The task involves careful comparison or trade-off analysis. Multi-criteria decisions, architectural choices, vendor evaluations.
- Your evaluation includes edge cases that lower-reasoning configurations miss. Measure those cases directly rather than inferring reasoning quality from fluent text.
- The task is consequential. Do not select more reasoning merely because the stakes are high. In finance, law, medicine, safety, or production operations, use authoritative sources, qualified review, validated controls, and a documented decision boundary. Then test whether the reasoning configuration improves the cases that matter.
Cases where a lower-reasoning baseline may be competitive:
- Latency-sensitive conversational chat. Use more reasoning only if its measured quality gain justifies the slower interaction.
- Generation and drafting where evaluations show no benefit. A lower-latency configuration may be better for creative iteration; compare output quality rather than assuming that more reasoning helps or hurts.
- Simple retrieval. A grounded lookup or lower-reasoning configuration may meet the quality target with less latency and cost. Verify the source either way.
- Rapid iterative refinement. Lower latency may improve the interaction, but compare final task success rather than counting turns alone.
- Tasks where you need auditable intermediate evidence. Ask for calculations, sources, test results, or other verifiable artefacts. A visible chain-of-thought narrative is not evidence that the conclusion is correct.
A useful decision rule is to increase reasoning only when the expected quality gain is worth the measured latency and cost for that workflow.
Prompt patterns to compare
Start with a direct, outcome-focused prompt, then compare these patterns when they could materially change the result:
1. Use a direct prompt as the OpenAI baseline
For OpenAI reasoning models, OpenAI’s reasoning best practices say that prompts such as “think step by step” are unnecessary and can sometimes hinder performance. Other providers expose different controls, so test a direct prompt against any more structured variant you are considering.
Bad: Think step by step. Solve this carefully. Show your work. [problem]
Good: [problem]
State the problem clearly and verify the conclusion. For other providers, follow current documentation and test the supported prompt or control.
2. Compare direct prompts with procedural scaffolding
Heavy scaffolding can duplicate work a reasoning model already performs. Start with the goal, relevant context, hard constraints, required evidence, and output format. Remove procedural steps only when an evaluation shows that the simpler prompt preserves the behavior you need.
Procedural candidate: First, list the key constraints. Then enumerate the options. Then evaluate each option against each constraint. Then pick. Then justify. Output format: …
Direct candidate: Help me decide between option A and option B. Context: […]
The simpler prompt gives the model room to choose an approach, while preserving the constraints and output contract you actually need.
3. Test technique stacking rather than assuming it helps
Chain-of-thought, self-critique, and tree-search prompts add instructions and tokens. Do not assume that stacking them improves the evaluated answer.
If your prompt includes “think step by step, then critique your own answer, then revise,” compare it with a simpler version. Keep the version that performs better on representative cases.
4. Use roles only when they carry task information
Elaborate personas often add unverifiable detail without changing the task. Use a role when it establishes scope, audience, policy, or terminology; otherwise prefer a direct description of the work. If the role appears to improve results, confirm that gain on the same evaluation cases rather than crediting the persona by intuition.
A short role can be useful for setting tone and register. Compare it with the equivalent concrete instructions when consistency matters.
Bad: You are a world-class senior backend engineer with 20+ years experience…
Good: Help me reason through this distributed-systems issue. [problem]
5. Ask for evidence, not hidden thinking
Some products hide or summarise internal reasoning. Asking the model to “show your reasoning” requests a generated explanation, not necessarily the model’s private trace or proof of correctness.
If you need auditability, ask for a concise rationale plus checkable evidence. Anthropic’s current API can return summarised or omitted thinking depending on the model and configuration; neither should replace verification of the result.
What to include in the prompt
A useful baseline includes:
Specifics. Provide the numbers, dates, exact constraints, files, and other inputs needed to solve and verify the task.
Open framing where the output contract permits it. “Here is the situation. Here is what I want to figure out. What do you think?” can be a useful candidate to compare with a more rigid template.
Honest uncertainty. Tell the model what you do not know. “I’m not sure if X or Y; help me identify what evidence would distinguish them” is more useful than hiding the ambiguity.
Permission to disagree. “Push back if my framing is wrong” or “tell me what I’m not considering” can expose alternatives that a confirmation-seeking prompt suppresses.
Concrete data. Spreadsheets, code, and documents give the model real artefacts to inspect. Share them only when the tool is approved for the data, and remove information the task does not need.
Worked examples
Example 1: A debugging task
Suppose you have a tricky bug.
Procedurally scaffolded prompt:
You are a senior software engineer specialising in TypeScript. Think step by step about this bug.
First, identify the relevant pieces of code. Second, trace through the data flow. Third, identify likely causes. Fourth, recommend a fix.
Here is the bug: [description] Here is the code: [code]
Direct prompt:
Help me find this bug.
Symptoms: [description] Relevant code: [code] What I’ve already tried: [list]
Compare the direct prompt with the scaffolded version on representative bugs, measuring correctness, required evidence, latency, and cost. Do not infer that the direct version is better from this example alone.
Example 2: A strategic decision
Structured prompt:
You are a senior strategy consultant. I’m trying to decide whether to launch product X. Apply the [framework name] framework. First, … [long structured prompt]
Direct prompt:
I’m trying to decide whether to launch product X. Context:
- We’re a 50-person company at $5M ARR.
- The product would take 2 quarters to build.
- It’s adjacent to but not directly competitive with our main product.
- Two of our top 10 customers have asked for it.
- Our team capacity is already strained.
Help me think through this. Push back on weak reasoning. Tell me what I’m not considering.
The direct prompt leaves room for the model to choose the analysis structure. Compare it with the structured variant against the same decision criteria before adopting either as a template.
Example 3: Complex code analysis
Procedurally scaffolded prompt:
Analyse this code for performance issues. Think step by step. First identify the data structures, then trace through the algorithm complexity, then point out specific bottlenecks. [code]
Direct prompt:
What’s slow about this code? It currently takes ~3 seconds on a typical input; I’d like it under 500ms.
[code]
Test whether the direct prompt identifies the relevant bottleneck and produces a valid fix. Verify every proposal with profiling data, tests, and benchmarks.

Pitfalls specific to reasoning models
A short list of things that catch even experienced users:
The latency. More reasoning can increase response time, sometimes materially. Measure the actual distribution for your model and workload, including timeouts, rather than planning from a generic estimate.
The cost. Providers account for reasoning differently, and model prices change. Measure billed input, output, and reasoning tokens on representative requests, then compare total task cost rather than assuming a fixed multiplier.
A response hits its generation limits. Reasoning and answer tokens share limits differently across providers. If a response stops early, inspect the provider’s stop reason and token accounting. Then narrow the task or adjust only the supported per-model controls; do not assume that every API accepts a generic thinking budget.
Long, stalled, or unproductive responses. You may observe unusually long latency, timeouts, repetition, or a final answer that misses the task. Record the observable failure and configuration. Then retry within your policy, narrow the task, or compare a different supported setting; do not claim to know the hidden cause from the output alone.
Fluent but unsupported conclusions. More reasoning does not prove correctness. On critical outputs, require sources, calculations, tests, and the assumptions or new evidence that would change the conclusion; self-reported confidence is not a calibrated guarantee.
Variable cost between requests. Reasoning-token use can vary with the task and configuration. Track actual usage rather than assuming each request costs the same or that cost scales predictably with subjective hardness.
A routed workflow to test
One candidate is a staged route:
- Lower-latency configuration to scope the task and identify the sub-questions.
- Higher-reasoning configuration for the evaluated sub-questions where the baseline misses requirements.
- Approved production configuration to format the verified result for its destination.
Compare this route with a single-configuration baseline. Measure task pass rate, evidence completeness, handoff errors, service-level compliance, end-to-end latency, and total cost. Extra stages are useful only when the final workflow performs better enough to justify their operational complexity.
A worked example: a market analysis task.
- Scoping stage: “I want to understand the market for X. Help me scope the analysis: what should I look at, what data do I need, what questions matter?”
- Analysis stage: “Given the verified data I’ve gathered, what does it imply for [specific strategic question]? Identify assumptions and evidence gaps.”
- Formatting stage: “Turn the verified findings into a one-page brief for our leadership team. Preserve the sources and uncertainty.”
This split is a testable workflow, not a guaranteed optimum. Measure it against a single-model baseline for quality, latency, and cost.
A few practical habits
Define the eligible baseline. The comparison must satisfy the same data, safety, tool, and output requirements as the reasoning candidate.
Predefine the decision rule. Choose the quality metric, service-level target, cost boundary, and minimum improvement before seeing the results.
Track your reasoning-model costs. Whether through subscription tier monitoring or API billing, get a feel for what your monthly reasoning-model bill looks like. Tune your usage accordingly.
Notice repeated failures in the lower-latency baseline. That is a useful trigger to test more reasoning on the same cases.
Compare one prompt change at a time. When a response is weak, test a shorter prompt, additional context, or a different supported reasoning setting separately so that the result is interpretable.
Choose with evidence
Reasoning configurations are not automatically better or worse. Start with a clear, direct prompt, retain real constraints and evidence requirements, and add structure or reasoning effort only when representative evaluations justify it.
Use reasoning deliberately and start with a simple prompt. A lower-latency model for exploration followed by more reasoning for evaluated bottlenecks is one pattern worth comparing with a single-model workflow.



