Skip to content

Prompt, Retrieve or Fine-Tune: How to Choose Where Model Behaviour Lives

Published: at 04:35 AMSuggest Changes

Three teams walked into the same quarterly review with three assistants built three different ways. One ran on a long, carefully maintained prompt. One retrieved from the document store at query time. One had been fine-tuned on last year’s policy pack. All three demoed well, and all three answered the same questions correctly.

The CIO asked a single follow-up. When the policy changes next quarter, who updates each of these, and how long does it take? The prompt owner answered in minutes. The retrieval owner answered in days, once the content team had republished and the index had refreshed. The fine-tuning team went quiet, because the honest answer was a data refresh, a training run, a fresh evaluation, a deployment and a governance review — and no one in the room could say who signs that off.

That silence is the real cost of technique selection, and it never shows up in a demo. Prompting, retrieval and fine-tuning get presented as three ways to get better answers. Treat them instead as three places to store behaviour, each with its own owner, change cycle, running cost and risk. Most teams pick the third because it was the technique in the last vendor demo they saw, and then spend a year discovering that behaviour has a maintenance bill wherever it lives.

Behaviour lives in one of three places

A prompt or context window changes what happens inside a single call. The model’s weights are untouched. Instructions, examples, output schemas and retrieved passages all live here, and the whole layer can be rewritten in one commit.

Retrieval changes what the model can see at the moment it answers. Facts arrive as tokens, so they can differ between users and between calls. What the model knows is governed by what you choose to put in front of it.

Fine-tuning changes the weights themselves. It alters how the model tends to behave, every time, for everyone using that model. That permanence is the whole appeal and the whole problem.

A vendor demo always looks impressive, because the fine-tuned model has been given a narrow task and a flattering prompt. What the demo never shows is the maintenance line: who re-tunes when the base model is upgraded, who owns the training data, and how you remove a fact that turns out to be wrong.

Six axes that actually decide it

Run a proposed behaviour through six questions before anyone commits engineering time.

How often does it change? This is the first filter and it settles most arguments. Microsoft Research tested how well fine-tuning injects new knowledge across five commercial models and found an average generalisation accuracy of 37% for learning new facts, dropping to 19% for updating facts the model already held. Two of the models could not learn or update knowledge through their fine-tuning interfaces at all. [Microsoft Research, December 2023] If a fact has a shelf life — a price, an entitlement, a policy clause, this quarter’s roadmap — the weights are the wrong home for it.

Who owns it? The layer you choose names the owner, whether or not you intended it to. A prompt belongs to the product owner. Retrieval belongs jointly to a content owner, who keeps the source truthful, and a platform team, who keeps the index and access controls working. Fine-tuned behaviour belongs to whoever runs the training pipeline, usually a specialist team far from the business process being served. If the person accountable for the outcome cannot change the behaviour, you have separated accountability from control — and the separation will surface as a delay, an exception or an incident.

What does evaluation cost? Prompt changes are cheap to test and cheap to roll back, so you can iterate quickly and measure. Retrieval requires you to measure two things separately: whether the right material was found, and whether the model used it well. Blur them and you will spend weeks tuning prompts against what is really a retrieval failure. Fine-tuning is the most expensive to evaluate, because you need a held-out set you did not train on, a stable task definition, and several hundred high-quality examples of the exact behaviour you want. Without that evidence you cannot tell whether the fine-tune helped, hurt or did nothing.

What does it cost to run? Retrieval adds a lookup and often a reranking step, so latency and per-call context tokens rise, and that bill scales with every request. Fine-tuning can move the other way: a specialised model can carry a shorter prompt, run on a smaller endpoint and cut unit cost at high volume. Microsoft’s agriculture study found the two gains stack rather than compete, with fine-tuning adding roughly six percentage points of accuracy and retrieval a further five on top. [Microsoft, January 2024] That result is worth holding onto. The techniques are complements, and the economics of each depend on your volume and latency budget, not on fashion.

What data-governance risk does it create? Retrieval enforces access control at query time. The model sees a document only if the person asking is entitled to it, and when the document is corrected or withdrawn, the change takes effect on the next query. Fine-tuning moves the data into the weights. You cannot enforce per-user access on knowledge that every call can reach, you cannot reliably delete a single fact, and you have to inspect the training set before it goes anywhere near a provider. That inspection obligation lands on your data protection and security leads, and it does not disappear because the data was already somewhere in your estate.

How reversible is it? A prompt is one commit. A retrieval index is a reindex. A fine-tuned model is a training run, a serving path and a re-certification, and the cost recurs every time the underlying base model changes. Put behaviour where you can change your mind quickly, and keep the slow, hard-to-reverse layer for the things you genuinely never want to move fast.

What should never go into the weights

Some material is simply wrong for fine-tuning, however well it tests.

Facts that change belong in context. Anything access-controlled belongs in retrieval, where entitlement is checked per request. Personal or sensitive data that a regulator, a customer or a data subject may ask you to delete belongs somewhere you can actually delete it.

There is a sharper warning. Fine-tuning a safety-aligned model can degrade its safety, and it does not take a malicious actor to do it. Researchers found that fine-tuning with a handful of adversarially chosen examples removed the guardrails from a commercial model for under a dollar, and that even benign, utility-focused datasets degraded alignment to a lesser degree. [ICLR, May 2024] Treat your model’s refusal and escalation behaviour as a control you must not casually overwrite. If a team’s fine-tune quietly softens how the assistant handles a sensitive request, you have altered a risk control without a review.

Combine them, in the right order

Retrieval for facts and fine-tuning for form is a sensible production pattern, and the evidence supports using both. Fine-tune a model to speak in your house style, to return a strict JSON shape, or to classify into a fixed taxonomy, and let retrieval supply everything that changes. The same split helps with small, stable corpora: Anthropic notes that a knowledge base under roughly 200,000 tokens may be simpler to place directly in the prompt with caching than to build retrieval infrastructure around. [Anthropic, September 2024]

The order matters more than the combination. Before any fine-tuning project, improve retrieval. Anthropic’s contextual-retrieval work showed that rewriting how chunks are built, adding keyword search alongside embeddings and reranking the results cut the top-20 retrieval failure rate from 5.7% to 1.9%, with the largest single gain coming from a change to chunk content rather than to any model. [Anthropic, September 2024] That is cheap engineering applied before anyone touches weights, and it frequently moves the metric the fine-tuning project was meant to fix.

Anthropic’s later guidance on context engineering draws the boundary clearly: treat the context window as a finite attention budget, keep the highest-signal tokens, and reserve the weights for stable behaviour. [Anthropic, September 2025] Facts in context, form in the weights, and a measurement before either.

The decision, in order

Work through these steps and stop at the first one that fits.

Can a clearer instruction or a better example fix it? Change the prompt. Is the corpus small and stable enough to sit in the prompt with caching? Do that and skip the retrieval infrastructure. Does the correct answer depend on facts your organisation owns or that change faster than a training cycle? Build retrieval. Is the failure one of format, tone or latency at high volume, with a measured baseline and a stable task definition? Fine-tune, and scope it to form rather than knowledge. Both? Combine them.

If you cannot describe what correct looks like well enough to write a hundred-question evaluation set, you are not ready for retrieval or fine-tuning. Build the evaluation first; every later decision depends on it.

The leadership test is a single question asked of every production assistant: name each behaviour it relies on, say where that behaviour lives, and name the person who changes it. If any answer is “in the model, and it would take a training run”, ask what happens when the underlying fact changes next quarter and who signs the change off. Put every behaviour at the lowest layer that can carry it, give it an owner, and make the change path measurable in days. Reserve the weights for the things you never want to change in a hurry, and treat anything else stored there as a maintenance liability you chose to create.

Sources


Previous Post
Who Pays When the Agent Is Wrong?
Next Post
Deflection Is Not Resolution: The Real Economics of AI Customer Service