2 system design questions lean on this idea. Each walks through the full answer.
A support team wants a chatbot that answers questions about their product. The product docs change almost every week as features ship. Someone on the team reaches for the most powerful tool they know: they collect a few thousand old support tickets, fine-tune a model on them, and deploy it.
It works. For about a week.
Then the pricing page changes, and the bot keeps quoting the old numbers with total confidence. A new feature launches, and the bot has never heard of it. Every time the docs move, the team has to gather new data, run another training job, evaluate it, and redeploy. They have turned a content problem into a machine learning problem, and now they are stuck retraining a model forever just to keep facts current.
The painful part is that they picked the wrong tool. The bot did not need new weights. It needed the current docs in front of it at the moment of the question. That is retrieval, not , and it would have cost them a fraction of the effort. Worse, the fine-tune actively hurt them in a way that is easy to miss at first: it welded stale facts into the model, so the bot did not just fail to know new things, it confidently asserted old ones. A model that says nothing is annoying. A model that insists last month's price is correct is a support ticket, a refund, and a trust problem all at once.
Most LLM projects that go sideways picked the wrong adaptation method, not the wrong model.
There are exactly three ways to make a general model do your specific job. Prompting, retrieval, and fine-tuning. Choosing well is one of the highest-impact decisions in an LLM system, because the choice sets your cost curve, your budget, your maintenance burden, and how fast you can react when the world changes. Get it right and the system is cheap to run and easy to keep honest. Get it wrong and you are paying to retrain a model on a treadmill, or watching a giant prompt re-bill on every call. This lesson is about how to make that call on purpose instead of by reflex.
Start with a general base model. It is smart and broadly capable, but it does not know your business and it does not follow your rules by default. You have three tools to fix that, and it is worth being precise about what each one is before you compare them.
Here they are side by side. For each, notice three things: what you actually change, how you steer it once it is live, and where it breaks down under pressure.

Notice the split. Two of the three never touch the model at all. That is not a minor detail. It is the whole game. When you only change what you send the model, a change ships in seconds and rolls back just as fast, because there is nothing to retrain and nothing to redeploy beyond a config. When you change the model itself, every change is a training run with its own data, evaluation, and release. Hold that distinction, because the next slide turns it into the single most useful mental model in the lesson.
Picture the same model as three layers you can influence. Prompting writes the top layer, RAG writes the middle one, writes the bottom one, and nothing reaches across. This is the picture to keep for the rest of the lesson.

Read it top to bottom.
| Approach | What it changes | What it does NOT change | How fast a change ships |
|---|---|---|---|
| Prompting | The instructions in the request | Weights, and any knowledge not in the prompt | Instant. Edit the string, redeploy |
| The facts placed in the context, per request | Weights, and how the model behaves | Instant. Update the index | |
| Fine-tuning | The weights themselves | Facts that were not in the training data | Slow. Retrain, evaluate, redeploy |
The most useful sentence in this whole lesson: Facts belong in what it sees, because facts move. Behavior belongs in what it is, but only once that behavior is stable enough to be worth baking in. Put a moving fact in the weights and you have signed up to retrain forever. Try to teach a genuinely new behavior through the prompt alone and you will chase reliability with ever longer instructions that still break on the hard cases. The layers tell you which lever can even reach the problem you have.
Here is how to choose without guessing. The first question does almost all the work, so ask it out loud: is the gap missing knowledge, or is it wrong behavior?
Those two problems pull in opposite directions, and almost every wrong choice in this space comes from misdiagnosing which one you have. Walk the tree.

Written as a function you could actually run, the same logic looks like this. Change the task attributes and watch the recommendation move.
Two rules fall out of this that catch most mistakes. First, never fine-tune to add facts that change. Second, never fine-tune before you have the labeled data to support it. If in doubt, prompt first. It is the cheapest experiment you will ever run, and it does something the harder tools cannot: it tells you what they would even need to fix. A day of prompt iteration will usually show you whether your real problem is knowledge or behavior, which is exactly the input this whole tree runs on.
The tree in the abstract is clean. Real projects are messier, so it helps to see the decision against the situations you will actually recognise. Find the row that sounds like your problem, and the recommended tool is usually the one whose failure mode you can most afford to live with.

The phrase "failure mode you can most afford" is doing real work there. Every approach fails, and choosing well is partly about choosing which failure you would rather debug at 2am. RAG fails by retrieving the wrong passage, which shows up as a confident answer grounded in the wrong document, and you fix it by improving retrieval. Prompting fails by drifting on edge cases the instructions did not anticipate, and you fix it by tightening the prompt or adding an example. fails by baking in something you did not mean to, and the fix is a whole retraining cycle. When two approaches would both work, prefer the one whose failure you can correct in minutes rather than weeks.
Because is the most common right answer and the most commonly botched, look at exactly what it does on a single request. The model is never modified at any point in this chain.

The chain is simple. Turn the question into a vector, search your document store for the closest chunks, paste those chunks into the prompt, and let the model answer from text it can see rather than from memory. Update the store and the very next request sees the new facts, with zero retraining. That is why the support team from the first slide should have reached here. When their pricing page changed, a RAG system would have needed nothing more than the new page in the index, and the next question would have quoted the right number.
The catch worth naming now: RAG is only as good as its retrieval. If the search returns the wrong chunks, the model gets bad context and answers wrong with total confidence. The prompt and the model can both be perfect and the answer will still be wrong, because the model faithfully summarized the wrong source. This is the failure mode that surprises teams, because it does not look like a retrieval bug from the outside, it looks like the model made something up. Retrieval quality, which means chunking, , and ranking, is where most of the real RAG engineering time goes, and it gets its own lessons later in this chapter. For now the point is narrower: RAG moves your hardest problem from the model to the search, which is usually a trade worth making because search is something you can measure and improve directly.
There is one axis where the three approaches separate harder than anywhere else, and it deserves its own picture: freshness. Freshness is not a quality score, it is a clock. It measures how long it takes, from the moment the world changes, until your model tells the truth about it.

Watch what happens when a single fact changes, say the pricing page updates this morning. With prompting, if the fact lives in the prompt, you edit a string and the change is live instantly. With RAG, you reindex the changed document and the next request sees it, usually within minutes and often automatically. With , the fact is welded into the weights, so reflecting the change means gathering data, retraining, evaluating, and redeploying, which is a project measured in weeks, and by the time it ships the next fact has probably already moved.
This is the clock that ruled fine-tuning out for the support team, before quality even entered the conversation. If your answers have to track a moving world, the freshness axis alone tells you the facts cannot live in the weights. It is also why the honest architecture for anything fact-heavy keeps the facts in a place you can edit cheaply, and reserves training for the things that genuinely do not change from one week to the next.
The choice is not only about quality and freshness. The three approaches have completely different cost, , and maintenance profiles, and at real traffic that difference decides the bill. Put them on every axis at once.

Read across any row and you see the trade you are actually making. Prompting is cheap to start and worst at hard behavior. RAG owns freshness and pays retrieval latency for it. gives the shortest prompt and the deepest behavior control, and cannot keep a single fact current. There is no column that wins every row, which is a hint about where this lesson ends up.
The cost story is the one people get most wrong, because the intuition points backwards. Prompting feels free, since there is no training run and no infrastructure. But a large few-shot prompt is re-sent and re-billed on every single call, so its cost is a straight line through the origin with a steep slope. A fine-tune front-loads a fixed training cost and then shortens every prompt, so its cost line starts high and climbs slowly. Plot both against volume and they cross.

You can reason about the crossover directly instead of guessing. Below is a tiny model of the three cost curves. Change the monthly volume and watch which approach is cheapest flip as traffic grows.
A few honest trade-offs to hold onto. Prompting is cheap to start but a giant prompt is re-sent on every call, so at high volume it can cost more than a fine-tune that shortened the prompt. RAG adds an call and a search to every request, which is real latency you have to budget for. Fine-tuning front-loads a large cost and heavy maintenance, and pays it back only when the behavior is stable and the traffic is high enough to amortize the training. Pick the profile that matches your reality, not the one that sounds the most advanced.
A chart with a crossing point is easy to nod at. Working it out once by hand shows you where the number comes from, and what it depends on. I will use only the prices written in the code above. They are illustrative prices, not quotes from any provider, so the method matters more than the result.
Each approach has a cost line of the form: fixed cost plus a price per request times the number of requests. Call the number of requests per month r.
Prompting against fine-tuning. They cost the same when 0.009 × r = 2,000 + 0.0009 × r. Take 0.0009 × r from both sides: 0.0081 × r = 2,000. So r = 2,000 ÷ 0.0081, which is about 246,900. That is the crossover near 247 thousand requests a month drawn on the chart.
Now bring RAG in. Prompting against RAG: 0.009 × r = 500 + 0.0045 × r, so 0.0045 × r = 500, and r = 500 ÷ 0.0045, about 111,100. RAG against fine-tuning: 500 + 0.0045 × r = 2,000 + 0.0009 × r, so 0.0036 × r = 1,500, and r = 1,500 ÷ 0.0036, about 416,700.
Check one point in the middle. At 250,000 requests, prompting costs 0.009 × 250,000 = 2,250 dollars. RAG costs 500 + 0.0045 × 250,000 = 500 + 1,125 = 1,625 dollars. Fine-tuning costs 2,000 + 0.0009 × 250,000 = 2,000 + 225 = 2,225 dollars. So at that volume RAG is the cheapest of the three, and the playground prints the same.
Here is the lesson inside the arithmetic. A cost chart only compares approaches that could each solve your problem. RAG is cheap at 250,000 requests, but if your gap is a strict output format, RAG does not fix it at any price. So first use the decision tree to find which approaches can close your gap. Only then compare their cost lines, and only between the ones still standing.
Before per-request economics ever matter, there is the up-front effort of standing each approach up at all, and that ladder never changes order.

Prompting is a same-day experiment with no data dependency. RAG is a short build that needs your documents and some retrieval plumbing. is a project with a data-collection problem sitting in front of it, because you cannot even start until you have thousands of clean labeled examples, and gathering those is often the longest part of the whole effort. That ordering is why the honest first move on almost any task is to prompt, measure, and only then reach for a heavier tool if the gap survives the cheap attempt.
Step back from cost and effort and look at raw capability, because it makes the case for combining tools impossible to miss. Line the capabilities you might want down one side and the three approaches across the top.

Read the grid by row, not by column. Freshness is a real strength only under . Deep, reliable behavior is a real strength only under fine-tuning. Cheap-to-start is a real strength only under prompting. A serious product usually wants more than one of those greens at the same time, and the grid makes the uncomfortable truth obvious: there is no single column you can pick to get them all. The only way to collect greens from different columns is to stop treating this as a choice between three tools and start treating it as a set of layers.
The framing of prompting versus RAG versus is useful for choosing, but real production systems often use two or three together, because the three solve different problems and the capability grid showed that no single one covers the field.
The pattern that shows up again and again: for the facts, a light fine-tune for the behavior, a cache in front for cost. The support bot keeps its knowledge current through retrieval, a small fine-tune fixes the one thing prompting could never make consistent, say a strict answer format or a specific brand voice, and a semantic cache means repeat questions never even reach the model.

The reason this works is that each tool does the job it is actually good at. Retrieval handles knowledge that moves. The fine-tune handles behavior that is stable and hard to describe in words. The cache handles cost and . Trying to force any single one of them to do all three jobs is exactly how projects end up brittle and expensive. Force fine-tuning to hold facts and you retrain forever. Force prompting to enforce a hard format and you fight edge cases with an ever growing instruction block. Force the model to answer every repeat question from scratch and you pay for the same tokens a thousand times. Split the work and every piece stays simple enough to reason about and to debug, which matters more than it sounds, because a system you can debug is a system you can keep running.
None of this is abstract. Teams shipping LLM features at scale make this exact call constantly, and the gap type decides it every time.

A company building a docs assistant, in the shape of a product like Notion, leans hard on retrieval, because the whole point is answering from your workspace, and your workspace changes every minute. a model on a workspace would be obsolete before the training job finished. That is a textbook knowledge gap, and RAG owns it.
A team building a coding assistant that must always emit a strict diff format or a specific function-calling schema has a behavior gap. No amount of retrieval fixes an output shape, because the facts are already correct, it is the structure that is wrong. If prompting cannot make the format reliable enough, a fine-tune on thousands of correctly formatted examples bakes the behavior in and shortens every prompt at the same time, which is the rare case where the heavier tool also makes each request cheaper.
And the most mature systems, like a support platform in the shape of a company such as Intercom or Zendesk, tend to combine them: retrieval to ground answers in the customer's own help center, and a fine-tune or careful prompting to enforce tone and structure. The lesson from all of them is the same one this whole track keeps repeating. The tool is not the hard part. Choosing the right tool for the actual gap is the hard part, and every one of these teams landed on their answer by asking the same first question: is the gap missing knowledge, or wrong behavior? Now you have a framework that turns that question into a decision.
4 questions - Score 80% to pass
A chatbot must answer from product docs that change every week. Which approach fits best?
What is the fundamental difference between fine-tuning and the other two approaches?
A team wants their model to always output a strict JSON schema, and prompting cannot make it reliable enough. They have 5,000 correctly formatted examples. Best move?
Your prompting prototype works, but at 1,000,000 requests a month the bill is alarming because the few-shot prompt is huge. The behavior is now stable. What is the reasonable next step?
There is a subtle corollary that trips teams up. Fine-tuning is excellent at shaping how a model responds and surprisingly bad at reliably teaching it new facts. People assume that if they train on documents, the model will memorize them. In practice it absorbs the style and the shape of that data far more reliably than the specific facts, and it will happily blend a half-remembered training fact with a confident guess. If you need a specific number to be correct, that number has to be in the context at answer time. The weights are for behavior, the context is for truth.