AI & Careers
RAG vs Fine-Tuning vs Prompt Engineering: Which Do You Need?
Quick answer: Start with prompt engineering. If the model still gets facts wrong or does not know your data, you have a knowledge problem and you need RAG. If it knows the facts but will not behave consistently, in your format or your tone, you have a behaviour problem and that is what fine-tuning is for. The single most expensive mistake in this area is reaching for fine-tuning to teach a model new information. Fine-tuning changes how a model behaves, not what it knows. OpenAI's own optimisation guide puts it plainly: "Prompt engineering is typically the best place to start," and you should "squeeze as much accuracy out of basic methods as you can before reaching for more complex RAG or fine-tuning."
The mistake almost everyone makes first
A team builds something on an LLM. It gets answers wrong about their own products, their own policies, their own data. Someone says the obvious-sounding thing: we should fine-tune it on our documents.
That instinct is wrong, and it is expensive. Fine-tuning is further training on additional data so the model "starts representing and mimicking the patterns and characteristics of the fine-tuning dataset" (Claude docs). It teaches patterns, style and format. It is a poor and costly way to teach facts, and the facts it does absorb are frozen at training time and cannot be updated without training again.
If your model does not know something, the fix is almost always to give it the information at the moment you ask. That is RAG.
What each one actually does
Prompt engineering
Changing what you ask and how you ask it: instructions, structure, examples, reasoning steps, and the settings that control variability. Nothing about the model changes. It is the cheapest and fastest of the three, it requires no infrastructure, and it is where every project should begin, because it is also how you find out what is actually broken.
RAG (Retrieval-Augmented Generation)
Fetching relevant information at query time and passing it to the model along with the question. Claude's documentation describes it as combining retrieval with generation "to improve the accuracy and relevance of the generated text, and to better ground the model's response in evidence", which lets the model "access and use information beyond its training data" (Claude docs). This is the answer to a knowledge gap, and it has a property fine-tuning does not: update the documents and the answers update immediately.
Fine-tuning
Further training the model on your examples so it internalises a pattern of behaviour. This is the right tool when you need a specific format or voice on every single response, and prompting gets you there only most of the time. It is the most expensive of the three in money, data and effort, and it is worth noting that Anthropic does not offer it for Claude at all: "The Claude API does not currently offer fine-tuning" (Claude docs).
The diagnostic that settles it
OpenAI's optimisation guide frames the choice as two different kinds of memory problem, and it is the most useful way to think about this. After you have a baseline from prompting, ask which of these you are looking at:
| Symptom | Problem type | Reach for |
|---|---|---|
| Gets facts about your business wrong | Context, or short-term memory | RAG |
| Does not know anything after its training cutoff | Context | RAG |
| Cannot cite where an answer came from | Context | RAG |
| Knows the facts, but the format drifts | Behaviour, or learned memory | Fine-tuning |
| Will not hold a consistent tone or voice | Behaviour | Fine-tuning |
| Follows instructions only most of the time | Behaviour | Fine-tuning |
| Vague, generic or unfocused answers | Instruction quality | Prompt engineering |
| Different answers to the same question | Sampling settings | Prompt engineering |
Read it twice, because it is the whole article: knowledge problems go to RAG, behaviour problems go to fine-tuning, and everything starts at prompting.
Side by side
| Dimension | Prompt engineering | RAG | Fine-tuning |
|---|---|---|---|
| What it changes | The request | The context supplied | The model itself |
| Adds new knowledge | Only what you paste in | Yes, this is its purpose | Poorly, and frozen at training time |
| Changes behaviour and format | Often enough | No | Yes, this is its purpose |
| Time to first result | Minutes | Days to weeks | Weeks, plus data preparation |
| Cost | Negligible | Moderate, ongoing | High up front, repeated on updates |
| Updating information | Immediate | Immediate, update the documents | Requires retraining |
| Can cite sources | No | Yes | No |
| Skills needed | None technical | Python, embeddings, evaluation | ML engineering and curated data |
| Available on Claude | Yes | Yes | Not via the API |
The order to actually work in
Write a decent prompt and measure it
Not a guess at a prompt: a proper one, with clear instructions, examples where they help, and a reasoning step for anything involving multi-step logic. Then build a small evaluation set so you can tell whether a change helped. Without that baseline you cannot diagnose anything, and you will not know whether RAG or fine-tuning improved matters or not.
Diagnose the failures, do not guess at them
Read the wrong answers. Are they wrong because the model lacked information, or because it had the information and behaved badly with it? That one question determines everything that follows.
Add retrieval if it is a knowledge problem
Most production failures are here. The model cannot know your contracts, tickets or product documentation unless you give them to it. This is also where most of the engineering work lives: chunking, retrieval quality, re-ranking and evaluation.
Consider fine-tuning only for stubborn behaviour
If after good prompting and good retrieval the remaining problem is that the output format or voice still drifts, fine-tuning is the right tool. Arriving here is much rarer than the internet suggests.
Can you combine them? Yes, carefully
These are not mutually exclusive, and serious systems often use prompting plus RAG together. Fine-tuning plus RAG is also possible, but it deserves a warning that comes from OpenAI's own worked example: in their Icelandic language case study, adding RAG to an already fine-tuned model reduced accuracy, from 87 to 83. Their conclusion is the lesson, which is to use "the right optimization tool for the right job" rather than stacking techniques and hoping.
More tooling is not more accuracy. Each addition brings its own failure modes, and retrieval that surfaces irrelevant context actively damages an answer the model would otherwise have produced correctly.
Where agents fit into this
A fourth option has become common enough to mention. Instead of retrieving a fixed set of documents, you can let the model decide what it needs, call a tool to fetch it, check the result and try again. That is an agentic system, and it is what people mean by autonomous retrieval or research agents.
It is more capable and considerably less predictable, and it belongs after you have retrieval working rather than instead of it. If your retrieval is unreliable, an agent calling it repeatedly produces unreliable results more expensively. Our Agentic AI course covers that layer, including self-healing retrieval and multi-agent supervision.
The part nobody budgets for: evaluation
Whichever route you take, the work that decides whether it succeeds is evaluation, and it is consistently the piece teams skip. Without a measurable way to tell whether an answer is good, you cannot know whether your new chunking strategy helped, whether a prompt change broke something else, or whether fine-tuning was worth the money.
This is also why "it works" is such a weak claim in AI engineering, and why the first thing a serious interviewer asks is how you knew your output was right. Build the evaluation set before you build the clever part.
Which should you learn first?
If you are deciding where to put your own study time rather than choosing a technique for a project, the order is the same as the engineering order, for the same reasons.
- Prompt engineering first. It needs no programming, it applies to every AI tool you already use, and it teaches you how models behave, which the other two assume. Our self-paced prompt engineering course covers the mechanism rather than a list of prompts to copy.
- RAG next, once you are comfortable in Python. It is the backbone of most applied AI work in production, and it is where you learn to measure quality properly. Our RAG course runs from embeddings through hybrid search, re-ranking and RAGAS evaluation.
- Agents after that. They extend retrieval with planning, memory and tools, and they assume you already understand what they are extending.
- Fine-tuning last, if at all. For most engineers it never becomes the daily job. Knowing precisely when it is the right answer is more valuable than being able to run one.
If you would rather have the full career path than individual skills, our Artificial Intelligence course and the Applied AI Forward Deployment Engineering Program cover this ground with live teaching and placement assistance.
Sources
- OpenAI, Optimizing LLM accuracy: the recommended order, the context-versus-behaviour diagnostic, and the Icelandic case study where RAG reduced the accuracy of a fine-tuned model from 87 to 83: developers.openai.com.
- Anthropic, Claude documentation glossary: definitions of fine-tuning, RAG, context windows, tokens and temperature, and the statement that the Claude API does not currently offer fine-tuning: platform.claude.com.
- Both sources were read directly at the time of writing. Model capabilities and vendor offerings in this area change quickly, so check the current documentation before making an architecture decision on the strength of any article, including this one. Written 3 October 2026.
FAQ
Frequently Asked Questions
What is the difference between RAG and fine-tuning?
RAG supplies information to the model at query time, so it can answer using data it was never trained on and you can update that data instantly. Fine-tuning further trains the model so it internalises patterns of behaviour, style and format. RAG changes what the model knows; fine-tuning changes how it behaves.
Should I fine-tune a model on my company documents?
Almost certainly not. Teaching facts through fine-tuning is expensive, imprecise, and freezes the information at training time so it cannot be updated without retraining. If the model needs to know your documents, retrieval is the correct approach, and it also lets the model cite where an answer came from.
Which should I try first?
Prompt engineering, always. OpenAI's own optimisation guide states that prompt engineering is typically the best place to start, and advises squeezing as much accuracy as possible out of basic methods before reaching for RAG or fine-tuning. It is also how you diagnose what is actually failing.
How do I know whether I have a knowledge problem or a behaviour problem?
Read the wrong answers. If the model is wrong because it lacked information it could not have known, that is a knowledge problem and RAG is the fix. If it had the information and still produced the wrong format, tone or structure, that is a behaviour problem and fine-tuning is the candidate.
Is fine-tuning available for Claude?
Not through the API. Anthropic's documentation states that the Claude API does not currently offer fine-tuning, and suggests contacting Anthropic if you want to explore it. For Claude, prompting, context management and retrieval carry the work that fine-tuning would do elsewhere.
Can I use RAG and fine-tuning together?
Yes, but carefully, and not automatically. In OpenAI's Icelandic case study, adding RAG to an already fine-tuned model reduced accuracy from 87 to 83. Stacking techniques adds failure modes, and retrieval that surfaces irrelevant context can damage an answer the model would otherwise have produced correctly.
Which is cheapest?
Prompt engineering by a wide margin, because it needs no infrastructure and no training run. RAG carries moderate ongoing cost in storage, retrieval and tokens. Fine-tuning is the most expensive, in money, in curated training data and in the effort to repeat it whenever requirements change.
Does RAG stop hallucination?
It reduces it but does not eliminate it. Grounding answers in retrieved evidence makes the model far less likely to invent facts, but if retrieval returns the wrong passages the model can still produce a confident wrong answer. The quality of retrieval determines the quality of the result, which is why evaluation matters.
What skills do I need for each?
Prompt engineering needs no programming at all. RAG needs working Python plus an understanding of embeddings, vector stores and evaluation. Fine-tuning needs machine learning engineering skills and, more importantly, the ability to build and curate a high-quality training dataset.
Where do AI agents fit in?
Agents are a layer above retrieval, not an alternative to it. Rather than fetching a fixed set of documents, the model decides what it needs, calls tools and can retry. Build agents after retrieval works, because an agent calling unreliable retrieval repeatedly just produces unreliable answers more expensively.
Why does the same prompt give different answers?
Because models sample from a probability distribution rather than selecting one fixed answer. Anthropic's documentation notes that even with temperature set to 0 results are not fully deterministic, and identical inputs may produce different outputs across API calls. Controlling this is part of prompt engineering.
What is the most common mistake teams make here?
Reaching for fine-tuning to solve a knowledge problem. It is intuitive, it is expensive, and it does not work well. The second most common is skipping evaluation, which leaves a team unable to tell whether any of their changes actually improved anything.
Do I need to learn fine-tuning to get an AI engineering job?
Knowing precisely when fine-tuning is the right answer matters more than being able to run one. Most applied AI roles spend their time on prompting, retrieval, evaluation and integration. Being able to explain why you did not fine-tune is often a stronger interview answer than having done it.
How long does each take to implement?
A meaningful prompt improvement takes minutes to hours. A working retrieval system takes days to weeks, depending on how messy the documents are. Fine-tuning takes weeks, with most of that spent assembling and cleaning training data rather than on the training itself.
Which should I learn first as a beginner?
Prompt engineering, because it needs no coding and teaches you how models behave, which the other approaches assume you understand. Then RAG once you are comfortable with Python, then agents. Fine-tuning last, and for many engineers it never becomes part of the daily job.
Keep Reading
Related Articles
AI & Careers
What Is MCP? Build Your First Model Context Protocol Server
A tested, end-to-end MCP tutorial: what the Model Context Protocol is, what changed in the 2026-07-28 spec, and how to build a server that exposes a database to Claude safely.
AI & Careers
FDE vs Solutions Engineer vs AI Engineer: The Real Differences
Forward Deployed Engineer, Solutions Engineer and AI Engineer compared: what each owns, how the interviews differ, how pay is structured, and which one fits you.
AI & Careers
Forward Deployed Engineer (FDE): The Complete Career Guide
What a Forward Deployed Engineer does, the skills and roadmap to become one, how FDE compares with ML and software engineering roles, and the India picture.