Llm Genai Ops

Fine-Tuning vs RAG vs Prompting: Choosing Your Approach

0 of 13 complete

0%

Contents

Back|Llm Genai OpsFine-Tuning vs RAG vs Prompting: Choosing Your Approach
1/13
36 min left
  1. Home
  2. AI Engineering: Evaluation, LLM Ops and Security
  3. LLM and GenAI Ops
  4. Fine-Tuning vs RAG vs Prompting: Choosing Your Approach
Prerequisites
Reading a Model Name: What 27B, FP8, A3B and Instruct Actually Meanrequired
Related Topics
Production RAG and LLM Systems: Running It, Not Just Building ItCore ConceptsWhy ML Models Fail in Production: The Production GapCore ConceptsModel Packaging and Containerization: Killing 'Works On My Machine' for MLCore ConceptsModel Serving and Inference APIs: Turning a Model File Into a ServiceCore ConceptsTraining Pipelines and Orchestration: Retraining That Runs ItselfCore Concepts
Previous lessonReading a Model Name: What 27B, FP8, A3B and Instruct Actually MeanNext lesson
1 of 13
Parameter-Efficient Fine-Tuning: LoRA and QLoRA

System Design

  • Foundation
  • Intermediate
  • Advanced
  • Capstone

AI Engineering

  • Foundation
  • Data, RAG and Agents
  • Evaluation, LLM Ops and Security

systemdesign.academy

  • Home
  • Glossary
  • Interview prep
  • Reviews
  • About
  • Privacy
  • Terms

Where this shows up in interviews

2 system design questions lean on this idea. Each walks through the full answer.

  • →Design ChatGPT
  • →Design RAG System

The Team That Fine-Tuned Its Way Into a Corner

A support team wants a chatbot that answers questions about their product. The product docs change almost every week as features ship. Someone on the team reaches for the most powerful tool they know: they collect a few thousand old support tickets, fine-tune a model on them, and deploy it.

It works. For about a week.

Then the pricing page changes, and the bot keeps quoting the old numbers with total confidence. A new feature launches, and the bot has never heard of it. Every time the docs move, the team has to gather new data, run another training job, evaluate it, and redeploy. They have turned a content problem into a machine learning problem, and now they are stuck retraining a model forever just to keep facts current.

The painful part is that they picked the wrong tool. The bot did not need new weights. It needed the current docs in front of it at the moment of the question. That is retrieval, not , and it would have cost them a fraction of the effort. Worse, the fine-tune actively hurt them in a way that is easy to miss at first: it welded stale facts into the model, so the bot did not just fail to know new things, it confidently asserted old ones. A model that says nothing is annoying. A model that insists last month's price is correct is a support ticket, a refund, and a trust problem all at once.

Most LLM projects that go sideways picked the wrong adaptation method, not the wrong model.

There are exactly three ways to make a general model do your specific job. Prompting, retrieval, and fine-tuning. Choosing well is one of the highest-impact decisions in an LLM system, because the choice sets your cost curve, your budget, your maintenance burden, and how fast you can react when the world changes. Get it right and the system is cheap to run and easy to keep honest. Get it wrong and you are paying to retrain a model on a treadmill, or watching a giant prompt re-bill on every call. This lesson is about how to make that call on purpose instead of by reflex.

Three Ways to Adapt a Model

Start with a general base model. It is smart and broadly capable, but it does not know your business and it does not follow your rules by default. You have three tools to fix that, and it is worth being precise about what each one is before you compare them.

  • . You write better instructions. Role, task, rules, and a few worked examples, all sent in the request. The model is untouched.
  • (RAG). You fetch relevant facts from your own data at request time and paste them into the prompt as . The model is still untouched. Only its context changes.
  • . You keep training the base model on your own labeled examples so a new behavior is baked into the weights. This is the only one that changes the model itself.

Here they are side by side. For each, notice three things: what you actually change, how you steer it once it is live, and where it breaks down under pressure.

Three ways to adapt one base model, shown as a comparison panel: a shared general base model with frozen weights forks into prompting in blue which writes better instructions and leaves the model untouched, RAG in emerald which fetches facts into the context and leaves the model untouched, and fine-tuning in violet which bakes behavior into the weights and is the only one that rewrites the model, each card listing what it changes, how you steer it, and where it breaks

Notice the split. Two of the three never touch the model at all. That is not a minor detail. It is the whole game. When you only change what you send the model, a change ships in seconds and rolls back just as fast, because there is nothing to retrain and nothing to redeploy beyond a config. When you change the model itself, every change is a training run with its own data, evaluation, and release. Hold that distinction, because the next slide turns it into the single most useful mental model in the lesson.

What Each One Actually Changes (and What It Does Not)

Picture the same model as three layers you can influence. Prompting writes the top layer, RAG writes the middle one, writes the bottom one, and nothing reaches across. This is the picture to keep for the rest of the lesson.

Model anatomy as three stacked layers with a lever for each: the top instructions layer in blue is written by prompting and ships in seconds, the middle context layer in emerald is written by RAG per request via an index update, and the bottom weights layer in violet is written only by a fine-tuning retrain cycle, with the rule that prompting and RAG change what the model sees while fine-tuning changes what the model is

Read it top to bottom.

ApproachWhat it changesWhat it does NOT changeHow fast a change ships
PromptingThe instructions in the requestWeights, and any knowledge not in the promptInstant. Edit the string, redeploy
The facts placed in the context, per requestWeights, and how the model behavesInstant. Update the index
Fine-tuningThe weights themselvesFacts that were not in the training dataSlow. Retrain, evaluate, redeploy

The most useful sentence in this whole lesson: Facts belong in what it sees, because facts move. Behavior belongs in what it is, but only once that behavior is stable enough to be worth baking in. Put a moving fact in the weights and you have signed up to retrain forever. Try to teach a genuinely new behavior through the prompt alone and you will chase reliability with ever longer instructions that still break on the hard cases. The layers tell you which lever can even reach the problem you have.

The Decision Framework

Here is how to choose without guessing. The first question does almost all the work, so ask it out loud: is the gap missing knowledge, or is it wrong behavior?

  • Knowledge gap: the model does not KNOW something. Your prices, your policies, last quarter's numbers, an internal runbook.
  • Behavior gap: the model knows plenty, but it will not ACT the way you need. Wrong format, wrong tone, ignores your output schema, rambles when you need three bullet points.

Those two problems pull in opposite directions, and almost every wrong choice in this space comes from misdiagnosing which one you have. Walk the tree.

Decision map for choosing an approach: the first question asks whether the gap is missing knowledge or wrong behavior, the knowledge branch asks whether facts change often or are too large for the prompt and routes yes to RAG and no to prompting, the behavior branch asks whether you have a thousand or more labeled examples with stable behavior and routes yes to fine-tuning and no to prompting first, with a hybrid recommendation of RAG for facts plus a light fine-tune for format underneath

Written as a function you could actually run, the same logic looks like this. Change the task attributes and watch the recommendation move.

Two rules fall out of this that catch most mistakes. First, never fine-tune to add facts that change. Second, never fine-tune before you have the labeled data to support it. If in doubt, prompt first. It is the cheapest experiment you will ever run, and it does something the harder tools cannot: it tells you what they would even need to fix. A day of prompt iteration will usually show you whether your real problem is knowledge or behavior, which is exactly the input this whole tree runs on.

When to Use Which, in Practice

The tree in the abstract is clean. Real projects are messier, so it helps to see the decision against the situations you will actually recognise. Find the row that sounds like your problem, and the recommended tool is usually the one whose failure mode you can most afford to live with.

When-to-use taxonomy table mapping situations to a recommended approach: docs that change weekly go to RAG, a small set of stable facts goes to prompting, a strict output format prompting cannot hold goes to fine-tuning, a specific brand voice at scale goes to fine-tuning, a huge private corpus that must cite sources goes to RAG, a brand-new feature with no labeled data goes to prompting first, and moving facts plus a format that must not move go to a hybrid of RAG and a light fine-tune

The phrase "failure mode you can most afford" is doing real work there. Every approach fails, and choosing well is partly about choosing which failure you would rather debug at 2am. RAG fails by retrieving the wrong passage, which shows up as a confident answer grounded in the wrong document, and you fix it by improving retrieval. Prompting fails by drifting on edge cases the instructions did not anticipate, and you fix it by tightening the prompt or adding an example. fails by baking in something you did not mean to, and the fix is a whole retraining cycle. When two approaches would both work, prefer the one whose failure you can correct in minutes rather than weeks.

How RAG Grounds an Answer

Because is the most common right answer and the most commonly botched, look at exactly what it does on a single request. The model is never modified at any point in this chain.

RAG request pipeline shown left to right: a user question is turned into a vector by an embedding step, a vector search finds the nearest chunks in your store, the top-k chunks are pulled out, a prompt is built from the question plus those retrieved facts, a frozen model reads that context, and a grounded answer comes out citing what it saw, with a note that updating the store makes the very next request see new facts with zero retraining

The chain is simple. Turn the question into a vector, search your document store for the closest chunks, paste those chunks into the prompt, and let the model answer from text it can see rather than from memory. Update the store and the very next request sees the new facts, with zero retraining. That is why the support team from the first slide should have reached here. When their pricing page changed, a RAG system would have needed nothing more than the new page in the index, and the next question would have quoted the right number.

The catch worth naming now: RAG is only as good as its retrieval. If the search returns the wrong chunks, the model gets bad context and answers wrong with total confidence. The prompt and the model can both be perfect and the answer will still be wrong, because the model faithfully summarized the wrong source. This is the failure mode that surprises teams, because it does not look like a retrieval bug from the outside, it looks like the model made something up. Retrieval quality, which means chunking, , and ranking, is where most of the real RAG engineering time goes, and it gets its own lessons later in this chapter. For now the point is narrower: RAG moves your hardest problem from the model to the search, which is usually a trade worth making because search is something you can measure and improve directly.

Freshness: Why Facts Do Not Belong in Weights

There is one axis where the three approaches separate harder than anywhere else, and it deserves its own picture: freshness. Freshness is not a quality score, it is a clock. It measures how long it takes, from the moment the world changes, until your model tells the truth about it.

Freshness propagation timeline for a single fact change such as a pricing update: prompting reflects it instantly by editing the string, RAG reflects it in minutes automatically by reindexing the doc, and fine-tuning takes weeks because it requires gathering data, training, evaluating, and redeploying, all drawn as bars along one shared time axis from now to weeks

Watch what happens when a single fact changes, say the pricing page updates this morning. With prompting, if the fact lives in the prompt, you edit a string and the change is live instantly. With RAG, you reindex the changed document and the next request sees it, usually within minutes and often automatically. With , the fact is welded into the weights, so reflecting the change means gathering data, retraining, evaluating, and redeploying, which is a project measured in weeks, and by the time it ships the next fact has probably already moved.

This is the clock that ruled fine-tuning out for the support team, before quality even entered the conversation. If your answers have to track a moving world, the freshness axis alone tells you the facts cannot live in the weights. It is also why the honest architecture for anything fact-heavy keeps the facts in a place you can edit cheaply, and reserves training for the things that genuinely do not change from one week to the next.

Cost, Latency, and Maintenance

The choice is not only about quality and freshness. The three approaches have completely different cost, , and maintenance profiles, and at real traffic that difference decides the bill. Put them on every axis at once.

Comparison matrix of prompting, RAG, and fine-tuning across seven properties: time to first version, data needed, per-request latency, keeps facts fresh, changes behavior, main ongoing work, and what each fails at, with colored pills showing that no single column wins every row

Read across any row and you see the trade you are actually making. Prompting is cheap to start and worst at hard behavior. RAG owns freshness and pays retrieval latency for it. gives the shortest prompt and the deepest behavior control, and cannot keep a single fact current. There is no column that wins every row, which is a hint about where this lesson ends up.

The cost story is the one people get most wrong, because the intuition points backwards. Prompting feels free, since there is no training run and no infrastructure. But a large few-shot prompt is re-sent and re-billed on every single call, so its cost is a straight line through the origin with a steep slope. A fine-tune front-loads a fixed training cost and then shortens every prompt, so its cost line starts high and climbs slowly. Plot both against volume and they cross.

Cost versus volume crossover chart with monthly cost on the y-axis and requests per month on the x-axis: prompting is a steep line through the origin, RAG has a small fixed cost and a medium slope, and fine-tuning starts with a high fixed training cost then climbs slowly, with a marked crossover near 247 thousand requests per month where the re-billed prompt overtakes the fine-tune that shortened it

You can reason about the crossover directly instead of guessing. Below is a tiny model of the three cost curves. Change the monthly volume and watch which approach is cheapest flip as traffic grows.

A few honest trade-offs to hold onto. Prompting is cheap to start but a giant prompt is re-sent on every call, so at high volume it can cost more than a fine-tune that shortened the prompt. RAG adds an call and a search to every request, which is real latency you have to budget for. Fine-tuning front-loads a large cost and heavy maintenance, and pays it back only when the behavior is stable and the traffic is high enough to amortize the training. Pick the profile that matches your reality, not the one that sounds the most advanced.

Find the Crossover by Hand

A chart with a crossing point is easy to nod at. Working it out once by hand shows you where the number comes from, and what it depends on. I will use only the prices written in the code above. They are illustrative prices, not quotes from any provider, so the method matters more than the result.

Each approach has a cost line of the form: fixed cost plus a price per request times the number of requests. Call the number of requests per month r.

  • Prompting: 0 + 0.009 × r. No fixed cost, a steep price per request, because the big prompt is sent every time.
  • : 500 + 0.0045 × r. A small fixed cost for the index and search, and a medium price per request.
  • : 2,000 + 0.0009 × r. A large fixed cost for training, then a short prompt and a low price per request.

Prompting against fine-tuning. They cost the same when 0.009 × r = 2,000 + 0.0009 × r. Take 0.0009 × r from both sides: 0.0081 × r = 2,000. So r = 2,000 ÷ 0.0081, which is about 246,900. That is the crossover near 247 thousand requests a month drawn on the chart.

Now bring RAG in. Prompting against RAG: 0.009 × r = 500 + 0.0045 × r, so 0.0045 × r = 500, and r = 500 ÷ 0.0045, about 111,100. RAG against fine-tuning: 500 + 0.0045 × r = 2,000 + 0.0009 × r, so 0.0036 × r = 1,500, and r = 1,500 ÷ 0.0036, about 416,700.

Check one point in the middle. At 250,000 requests, prompting costs 0.009 × 250,000 = 2,250 dollars. RAG costs 500 + 0.0045 × 250,000 = 500 + 1,125 = 1,625 dollars. Fine-tuning costs 2,000 + 0.0009 × 250,000 = 2,000 + 225 = 2,225 dollars. So at that volume RAG is the cheapest of the three, and the playground prints the same.

Here is the lesson inside the arithmetic. A cost chart only compares approaches that could each solve your problem. RAG is cheap at 250,000 requests, but if your gap is a strict output format, RAG does not fix it at any price. So first use the decision tree to find which approaches can close your gap. Only then compare their cost lines, and only between the ones still standing.

What Each Approach Costs You to Stand Up

Before per-request economics ever matter, there is the up-front effort of standing each approach up at all, and that ladder never changes order.

Effort-to-ship bar chart showing time to a first working version: prompting takes hours with no data and ongoing prompt edits, RAG takes days needing your documents with ongoing retrieval-quality work, and fine-tuning takes weeks needing thousands of labeled pairs with ongoing retraining and evaluation

Prompting is a same-day experiment with no data dependency. RAG is a short build that needs your documents and some retrieval plumbing. is a project with a data-collection problem sitting in front of it, because you cannot even start until you have thousands of clean labeled examples, and gathering those is often the longest part of the whole effort. That ordering is why the honest first move on almost any task is to prompt, measure, and only then reach for a heavier tool if the gap survives the cheap attempt.

Step back from cost and effort and look at raw capability, because it makes the case for combining tools impossible to miss. Line the capabilities you might want down one side and the three approaches across the top.

Capability heatmap with capabilities down the side and prompting, RAG, and fine-tuning across the top, cells colored green for a real yes, amber for partial, and red for no: fresh facts are green only under RAG, custom behavior is green only under fine-tuning, cheap-to-start is green only under prompting, and no single column is green all the way down

Read the grid by row, not by column. Freshness is a real strength only under . Deep, reliable behavior is a real strength only under fine-tuning. Cheap-to-start is a real strength only under prompting. A serious product usually wants more than one of those greens at the same time, and the grid makes the uncomfortable truth obvious: there is no single column you can pick to get them all. The only way to collect greens from different columns is to stop treating this as a choice between three tools and start treating it as a set of layers.

You Do Not Have to Pick Just One

The framing of prompting versus RAG versus is useful for choosing, but real production systems often use two or three together, because the three solve different problems and the capability grid showed that no single one covers the field.

The pattern that shows up again and again: for the facts, a light fine-tune for the behavior, a cache in front for cost. The support bot keeps its knowledge current through retrieval, a small fine-tune fixes the one thing prompting could never make consistent, say a strict answer format or a specific brand voice, and a semantic cache means repeat questions never even reach the model.

Hybrid architecture with three labeled layers: a semantic cache handles cost and latency by returning repeat questions immediately, on a miss a retriever plus store handles knowledge that moves by grounding the answer in current facts, then a fine-tuned model handles stable behavior by enforcing the strict format or brand voice, producing an answer that is fresh, correctly shaped, and cheap at scale

The reason this works is that each tool does the job it is actually good at. Retrieval handles knowledge that moves. The fine-tune handles behavior that is stable and hard to describe in words. The cache handles cost and . Trying to force any single one of them to do all three jobs is exactly how projects end up brittle and expensive. Force fine-tuning to hold facts and you retrain forever. Force prompting to enforce a hard format and you fight edge cases with an ever growing instruction block. Force the model to answer every repeat question from scratch and you pay for the same tokens a thousand times. Split the work and every piece stays simple enough to reason about and to debug, which matters more than it sounds, because a system you can debug is a system you can keep running.

How This Plays Out in Practice

None of this is abstract. Teams shipping LLM features at scale make this exact call constantly, and the gap type decides it every time.

Real-world case panel with three examples: a docs assistant in the shape of a product like Notion has a knowledge gap because the workspace changes constantly so it uses RAG, a coding assistant that must emit a strict diff or schema has a behavior gap so it uses fine-tuning, and a mature support platform in the shape of Intercom or Zendesk has both gaps so it uses a hybrid of retrieval for grounding and a fine-tune or prompting for tone

A company building a docs assistant, in the shape of a product like Notion, leans hard on retrieval, because the whole point is answering from your workspace, and your workspace changes every minute. a model on a workspace would be obsolete before the training job finished. That is a textbook knowledge gap, and RAG owns it.

A team building a coding assistant that must always emit a strict diff format or a specific function-calling schema has a behavior gap. No amount of retrieval fixes an output shape, because the facts are already correct, it is the structure that is wrong. If prompting cannot make the format reliable enough, a fine-tune on thousands of correctly formatted examples bakes the behavior in and shortens every prompt at the same time, which is the rare case where the heavier tool also makes each request cheaper.

And the most mature systems, like a support platform in the shape of a company such as Intercom or Zendesk, tend to combine them: retrieval to ground answers in the customer's own help center, and a fine-tune or careful prompting to enforce tone and structure. The lesson from all of them is the same one this whole track keeps repeating. The tool is not the hard part. Choosing the right tool for the actual gap is the hard part, and every one of these teams landed on their answer by asking the same first question: is the gap missing knowledge, or wrong behavior? Now you have a framework that turns that question into a decision.

Knowledge Check

Knowledge Check

4 questions - Score 80% to pass

Q1

A chatbot must answer from product docs that change every week. Which approach fits best?

Q2

What is the fundamental difference between fine-tuning and the other two approaches?

Q3

A team wants their model to always output a strict JSON schema, and prompting cannot make it reliable enough. They have 5,000 correctly formatted examples. Best move?

Q4

Your prompting prototype works, but at 1,000,000 requests a month the bill is alarming because the few-shot prompt is huge. The behavior is now stable. What is the reasonable next step?

prompting and RAG change what the model sees, fine-tuning changes what the model is.

There is a subtle corollary that trips teams up. Fine-tuning is excellent at shaping how a model responds and surprisingly bad at reliably teaching it new facts. People assume that if they train on documents, the model will memorize them. In practice it absorbs the style and the shape of that data far more reliably than the specific facts, and it will happily blend a half-remembered training fact with a confident guess. If you need a specific number to be correct, that number has to be in the context at answer time. The weights are for behavior, the context is for truth.