Choosing between prompting, RAG, and fine-tuning (and when to combine)
Advanced12 min readAI for Business

Choosing between prompting, RAG, and fine-tuning (and when to combine)

Prompting, RAG, and fine-tuning are the three big levers for adapting LLMs to your problem. Each is right for some problems and wrong for others. A framework for choosing, the realistic costs of each, and the production patterns where combining them shines.

What you should be able to do

Prompting changes what the model is asked. RAG supplies retrievable evidence at inference time. Fine-tuning adapts model behavior on a measured task; it is not a maintainable factual database. Choose from the failure you can demonstrate, and combine methods only when evaluation justifies the added system.

Saved only in this browser.
In this article

A team gets the brief: “Make our AI better at handling our specific use case.” They have a few choices on how to do it. They can write better prompts. They can build a RAG system to feed the model relevant data. They can fine-tune the model on examples of their domain.

These are not interchangeable. They solve different problems. Done wrong, you spend three months and a budget fine-tuning when the answer was a better prompt. Or building elaborate RAG infrastructure when fine-tuning would have been simpler. Or stuck on prompts when the model fundamentally can’t do what you need.

Here is a framework for choosing and, more importantly, for running a fair comparison. It does not publish universal project costs or accuracy gains. Primary references include the original RAG paper, OpenAI’s current fine-tuning guide, and Hugging Face PEFT documentation.

What each actually does

A clean distinction:

Prompting changes what the model is asked. You give the model better instructions, examples, format requirements, context. The model itself doesn’t change; the input does.

RAG changes what data the model sees. You retrieve relevant information at query time and include it in the prompt. The model has fresh, specific, dynamic data without being trained on it.

Fine-tuning changes model weights to adapt behavior on a measured task. It may affect style, format, task performance, and memorized associations. It is not a reliable or maintainable way to install a current factual database.

These solve different problems:

  • Instructional gap: the model could do the task if asked correctly. → Prompting.
  • Knowledge gap: the model needs information it doesn’t have. → RAG.
  • Capability gap: the model can’t reliably do the task even with good prompts and context. → Fine-tuning.

Knowing which gap you have is half the battle.

Diagnosing the gap

When the AI isn’t doing what you need, ask:

Could a smart person, given just the prompt, do this task?

If yes → instructional gap. Better prompt should fix it.

If no, could they do it if you gave them relevant reference material?

If yes → knowledge gap. RAG can fix it.

If no, could they do it after extensive practice and feedback?

If yes → capability gap. Fine-tuning might fix it.

If no → maybe the task isn’t solvable by an LLM. Reconsider the problem.

Do not assume how common each gap is. Label failures from a representative evaluation set: instruction-following, missing/incorrect context, retrieval failure, capability failure, policy failure, or downstream integration failure. That distribution tells you where to invest.

Prompting: the underrated lever

Prompting is the cheapest, fastest, and most often the right answer. Yet teams skip past it to RAG or fine-tuning.

A few things you can do with prompting alone:

  • Change tone, format, length.
  • Apply reasoning patterns (chain-of-thought, self-critique).
  • Encode constraints (do X, don’t do Y).
  • Encode policies and guardrails.
  • Adapt to specific use cases (different prompts for different features).
  • Improve consistency through few-shot examples.

What you can’t do with prompting alone:

  • Get the model to know facts it doesn’t.
  • Make a small model behave like a large model.
  • Fundamentally change the model’s voice or style on a deep level.
  • Speed up the model’s inference.

A reasonable first experiment is the lowest-cost prompt baseline. Time-box it by number of evaluated variants and a stopping rule, not by an arbitrary week. Move on when additional prompt changes no longer improve the held-out metric or violate another requirement.

Prompt engineering effort

Run prompt variants against the same train/development split and preserve a held-out test set. Change one design dimension at a time—task contract, examples, context, output schema, or tool set—and record quality, latency, and cost. Stop when the predeclared acceptance bar is met or the next change fails to improve the development set. A calendar duration cannot prove that prompting is exhausted.

What good prompts look like

Candidate prompt elements to test include:

  • Clear role and task.
  • Specific format requirements.
  • Representative examples when an eval shows they help.
  • Explicit constraints (what to do, what not to do).
  • Edge case handling.
  • Output schema.

Use the shortest prompt that preserves the measured behavior and necessary constraints. Paragraph count is not a quality metric.

RAG: the knowledge fix

RAG is a candidate when:

  • The model needs factual information it doesn’t have.
  • The information changes (live data, recent events, account-specific data).
  • The information is specific to your domain or organization.
  • You need citations / provable grounding.

It’s the wrong tool when:

  • The problem is instructional, not knowledge.
  • The data is small enough to fit in a prompt directly.
  • You need the model to do something differently, not just know something different.

The realistic cost

A RAG system is real engineering:

  • Build: ingestion, permissions, chunking, indexing, retrieval, evaluation, and operations; estimate from your sources and acceptance criteria.
  • Operate: ongoing — keeping the index updated, monitoring quality, fixing issues.
  • Infrastructure: parsing/OCR, storage, embedding, retrieval, reranking, generation, backups, and telemetry; price them from dated quotes and measured volumes.
  • Per-query cost: may add query embedding, retrieval, reranking, and context tokens, but can also reduce wasted context or enable a cheaper generator. Measure the complete path.

Approve the investment only when the evaluated retrieval path improves the defined outcome enough to justify ingestion, permission, freshness, latency, cost, and operational work.

RAG quality is a journey

There is no transferable “week-one quality” percentage. Baseline lexical, vector, and hybrid retrieval on a labeled set; report recall and final-answer metrics by slice; then change one stage at a time. Ship only after the business-specific error and permission checks pass.

Fine-tuning: when prompting and RAG aren’t enough

Fine-tuning is the right tool when:

  • You have a clear capability gap — the model can’t reliably do the task even with good prompts and context.
  • You have enough representative, licensed, privacy-reviewed training data to improve held-out performance. Determine sufficiency with a learning curve, not a universal example count.
  • You need consistent, narrow behavior (a specific style, a specific output format, a specific domain).
  • Inference cost / latency matters (a fine-tuned smaller model can be cheaper than a generic larger one).

It’s the wrong tool when:

  • Your data is in flux (the fine-tuned model will be stale fast).
  • You have not built the relevant prompt, constrained-output, retrieval, or tool baseline for comparison.
  • You don’t have good evals (you can’t tell if fine-tuning helped).
  • The task needs very current information (fine-tuning is a snapshot).
  • The task primarily needs current, attributable facts; retrieval or authoritative tools are easier to update and cite than weight changes, although they still require evaluation.

Types of fine-tuning

Full fine-tuning: all model weights are eligible for update. It has the largest trainable surface and compute/storage burden, but does not inherently produce the best held-out result.

LoRA (Low-Rank Adaptation): trains low-rank adapters while the base weights remain frozen, reducing trainable parameters and commonly reducing optimizer-memory needs. Compare quality and serving support with full tuning for the actual task.

QLoRA: backpropagates through a frozen quantized base model into LoRA adapters. It can reduce memory requirements; quality is empirical, not a fixed “lower at scale” rule.

Prompt tuning / prefix tuning: trains continuous prompt or prefix parameters. Hardware, provider support, task quality, and serving integration determine whether the smaller trainable surface is useful.

Instruction tuning: supervised adaptation on instruction-response examples. It is common in model development and can also be an application-team experiment when the model licence, data, infrastructure, and evals support it.

RLHF / DPO / KTO: different preference-optimization families; they are not interchangeable and do not all consume identical data. Use the pinned implementation’s primary paper and documentation, then test reward/preference quality and safety regressions.

Parameter-efficient tuning is a useful candidate when full tuning does not fit the hardware or iteration budget. This article has no representative survey proving what most teams use or a universal best balance.

The realistic cost

Fine-tuning cost and duration depend on base model, sequence lengths, tokens, epochs, hardware, distributed strategy, failures, and evaluation. Build the estimate from a small instrumented run:

  • Data prep: provenance, permission, deduplication, privacy review, formatting, and quality labeling.
  • Training: measured tokens per second × planned tokens, plus checkpointing, failed runs, and storage.
  • Evaluation: base versus tuned model on held-out target, safety, and general-capability sets.
  • Iteration: budget only after defining what result would justify another run.
  • Deployment: provider availability, adapter serving, rollback, monitoring, and retention obligations.
  • Maintenance: retraining when data updates, when the base model updates, when the use case shifts.

Include engineering and reviewer time; compute alone is not total cost. Proceed only when the measured improvement is worth the continuing serving and retraining burden.

When fine-tuning shines

Scenarios where a fine-tuning experiment may be justified:

Stable task behavior. Fine-tuning may improve format or style on a particular distribution, but constrained decoding can already guarantee supported schema shapes. Compare prompt-only, constrained-output, and tuned baselines rather than promising 95% versus 99%.

Specialized language or syntax. Medical terminology, legal phrasing, or an internal DSL may improve on held-out examples, but high-stakes correctness still needs external validation, current evidence, and qualified review.

Style/voice. A tuned model may improve a scored style rubric on the target distribution. It does not “bake in” perfect consistency; measure content quality, safety, and drift as well as voice.

Latency/cost experiment. A smaller tuned model may clear the same acceptance bar with lower measured serving cost or latency. Include training, evaluation, capacity, reliability, and maintenance before claiming a payoff.

Behavioral adaptation. Safety tuning may improve a measured refusal behavior, but it can also over-refuse, under-refuse, or regress other capabilities. It must remain one layer beside deterministic authorization, policy enforcement, monitoring, and human gates.

When fine-tuning fails

Common ways fine-tuning disappoints:

Insufficient or unrepresentative data. Small, narrow datasets sometimes help and large noisy datasets sometimes hurt. Plot held-out performance as training data increases and inspect slice coverage before buying more compute.

Bad data. Garbage in, garbage out. Inconsistent, low-quality examples produce inconsistent, low-quality models.

Catastrophic forgetting. Heavy fine-tuning on narrow tasks can hurt general capabilities. The model gets good at your task but worse at everything else.

Stale knowledge. Fine-tuned model is a snapshot. New information requires retraining. For dynamic domains, this is a perpetual cost.

Base model improvements outpace fine-tuning. The base model improved enough that the fine-tune is no longer better. You’re now maintaining a fine-tune of an outdated base.

Evaluation problems. Without solid evals, you don’t know if fine-tuning helped, hurt, or had no effect. Many “successful” fine-tunes are placebo wins.

The combination patterns

Some production systems combine these methods, but every added method expands the evaluation and operating surface.

Combination 1: Prompted RAG

A candidate for applications that need dynamic, attributable knowledge.

  • Carefully designed prompts encode instructions, format, constraints.
  • RAG provides current, specific information.
  • No fine-tuning; rely on strong base model.

Approve it only if retrieval improves the defined task and permission, freshness, citation, latency, cost, and failure tests pass.

Combination 2: Fine-tuned model + RAG

When you need both behavioral specialization and dynamic knowledge.

  • Fine-tune for voice, format, domain.
  • RAG for current information.
  • Prompts orchestrate.

Example: a fine-tuned model for a specific company’s customer support voice, with RAG over current policies and documentation. The fine-tune handles the consistent voice; RAG handles the changing knowledge.

Combination 3: Specialized fine-tunes for specific tasks

Different fine-tunes for different parts of a system.

  • Classification fine-tune for routing.
  • Summarization fine-tune for digests.
  • Generation fine-tune for customer responses.
  • Each smaller, faster, specialized.

Used when scale and cost optimization matter. Each fine-tune does its narrow job well; orchestration calls them.

Combination 4: Fine-tuned router + general models

The router is fine-tuned to classify queries reliably. Once classified, queries go to general models for the actual work.

The fine-tune is small, fast, narrow. The expensive general work is done by general models, kept current.

This combines economy (fine-tune is small) with capability (general models for the hard work).

A reference folder and instruction card support an adjustable mechanism
AI-generated illustration of combining prompting, retrieval, and model adjustment in one bounded workflow.

The decision framework

A practical decision flow:

Question 1: Is the problem solvable with the current model and a good prompt?

If yes: build and evaluate a prompt baseline. Ship only if the acceptance and safety criteria pass.

If no, go to Question 2.

Question 2: Does the problem involve knowledge the model doesn’t have?

If yes: compare direct context, lexical retrieval, vector retrieval, and hybrid retrieval as appropriate. Estimate the implementation after source and permission discovery.

If no, go to Question 3.

Question 3: Is the problem about consistent format, narrow domain, or specific behavior?

If yes, and you have representative data with the necessary rights and privacy controls: run a small fine-tuning experiment and compare it with the unchanged baseline.

If you don’t have the examples: invest in collecting them, OR try better prompting / RAG further before fine-tuning.

Question 4: Have you done the eval work to know which approach actually helps?

This question applies at every step. Without evals, you’re guessing.

Production examples

Illustrative combinations (not measured case studies):

Example 1: AI customer support

Setup: A SaaS company’s customer support AI handles tier 1 inquiries.

Components:

  • Strong prompts for tone, format, escalation policies.
  • RAG over current docs, policies, ticket history.
  • Lightweight fine-tune on the company’s specific voice and escalation patterns (1,500 examples curated from past tickets).

Outcome (illustrative scenario): Handles a large share of tier-1 tickets autonomously when measured against that team’s eval set. The fine-tune accounts for the consistent voice; RAG keeps it accurate; prompts handle the policies.

Setup: A legal-tech product reviews contracts for risks.

Components:

  • Detailed prompts encoding what to look for (legal categories, severity rubric).
  • RAG over relevant case law and precedent.
  • No fine-tuning; reasoning models handle the heavy lifting.

Decision criterion: compare clause-level recall, false reassurance, citation correctness, jurisdiction coverage, and qualified-lawyer review. No method is approved for legal reliance merely because a base model has seen legal text.

Example 3: Code completion in a custom DSL

Setup: A specialized data tool with its own DSL.

Components:

  • Prompts with examples.
  • No RAG (the DSL is small enough to fit in context).
  • A rights-cleared training set whose size is chosen from coverage and a learning curve rather than a copied example count.

Decision criterion: compare parse success, semantic correctness, and execution safety for prompted and tuned versions on held-out programs. The illustrative setup does not establish that fine-tuning is essential for a real DSL.

Example 4: Internal company assistant

Setup: A general assistant for company employees.

Components:

  • Strong system prompts (voice, behavior, refusals).
  • RAG over company wiki, Slack, docs.
  • No fine-tuning; the company’s “voice” is captured in prompts.

Decision criterion: compare prompt-only and retrieval variants on current-answer correctness, citations, permissions, abstention, latency, cost, and reviewed style. Fine-tune only if a remaining, valuable behavioral gap is demonstrated.

Mistakes we see

A few patterns of misallocation:

Mistake 1: Reaching for fine-tuning first. Teams hear “we should fine-tune our own model” and start there. Test prompting and retrieval baselines first; many apparent fine-tuning problems are instruction or missing-context problems, and the baseline gives you evidence for the decision.

Mistake 2: Skipping RAG when it’s the answer. Teams build elaborate prompts to “remind” the model of company info that should obviously be retrieved at query time. Better to just retrieve.

Mistake 3: Fine-tuning without evals. “We fine-tuned and it’s better now.” Without a held-out baseline and slice results, the change cannot be attributed and regressions remain hidden.

Mistake 4: Stale fine-tunes. Base models, serving stacks, data, and product requirements change. Re-run the unchanged baseline and acceptance suite before retraining or continuing to serve an adapter.

Mistake 5: Treating weights as a current database. Training may create memorized associations without update, provenance, or deletion guarantees. Keep authoritative facts in governed sources or tools and test retrieval; use tuning for a demonstrated task behavior, not as the system of record.

Mistake 6: No prompt stopping rule. Iterating without a held-out set can overfit examples indefinitely. Stop or escalate when predeclared metrics plateau.

Mistake 7: Over-engineering RAG when prompting could do it. A 50K-token company doc dumped in the prompt is sometimes simpler than RAG. Especially for small corpora.

Cost and effort comparison

A comparison worksheet—fill it with measured or quoted values for the same acceptance bar:

ApproachEffortCost (one-time)Cost (per query)Maintenance
Promptingvariants × eval runtime + reviewmodel calls + reviewer timemeasured tokens/toolsprompt/model/eval updates
RAGsource discovery + ACL + pipeline + evalingestion, index, engineeringretrieval + rerank + generationsource sync, permissions, eval
Fine-tuning (LoRA)data + training + eval + servingrights review, compute, engineeringserving hardware/providerdata, base model, safety eval
Prompting + RAGcombined critical pathcombined, less shared workmeasured full traceprompt, sources, retrieval, eval
All threecombined critical pathcombined, less shared workroute-specifichighest operational surface

The right choice depends on the measured gap and the complete operating model. Prompting plus retrieval is common for dynamic knowledge, but it is not a universal sweet spot.

Pick the gap, then the lever

Prompting, RAG, and fine-tuning solve different problems. Picking right requires clear diagnosis: is this an instructional gap, a knowledge gap, or a capability gap?

The order to try them:

  1. Build the simplest relevant baseline. Often this is prompting or constrained output, but a retrieval or tool baseline may be simpler when facts are already in an authoritative system.
  2. Add retrieval for a demonstrated evidence/freshness gap, then evaluate the complete ingestion, permission, retrieval, and generation path.
  3. Run a fine-tuning experiment for a demonstrated behavioral or task-performance gap when rights-cleared representative data and serving support exist.
  4. Combine methods only when each component produces a measured incremental gain worth its operating surface.

Without held-out evaluations, you cannot attribute which approach helped. Record the gap, baseline, acceptance criteria, adverse slices, complete cost, and rollback before choosing the lever.

Read next

Continue through the same learning path with the next practical articles.