RAG vs fine-tuning comes down to what is broken. If the model does not know your facts (your policies, your catalog, last week's prices), use RAG. If it knows enough but behaves inconsistently (wrong format, wrong tone, shaky on one narrow task), consider fine-tuning. And before either, use prompt engineering, because OpenAI's own guide calls it "typically the best place to start" and says it is often the only method you need. The order for a product manager is: prompt, then retrieve, then train, with evals deciding each step.
AllthingsPM is an AI PM course and PM interview prep platform. Its course was built from 604 real PM job postings and teaches this exact decision in the Agents and agentic architecture chapter, with a graded case study at the end, so you can practise the call instead of just reading about it.
RAG vs fine-tuning vs prompt engineering: what is the difference?
All three change what a large language model produces. They differ in where the change lives.
- Prompt engineering changes the instructions. You rewrite the system prompt, add examples, set a format. Nothing about the model changes, and you can ship a new version in minutes.
- RAG (retrieval-augmented generation) changes the context. Before the model answers, your system fetches relevant documents and pastes them into the prompt. OpenAI describes it as "the process of Retrieving content to Augment your LLM's prompt before Generating an answer." The idea comes from a 2020 paper by Patrick Lewis and colleagues.
- Fine-tuning changes the model's weights. You continue training on a smaller, task-specific dataset, so the new behaviour is baked in rather than asked for each time.
OpenAI frames this as two axes. Context optimization fixes knowledge that is missing, outdated or proprietary. LLM optimization fixes inconsistent results, formatting and reasoning. RAG sits on the first axis, fine-tuning on the second, and prompting touches both.
| Prompt engineering | RAG | Fine-tuning | |
|---|---|---|---|
| Where AllthingsPM teaches it | Make the API call yourself and Context windows | Context as a budget: when RAG is the fix | Pretraining vs post-training and the fix hierarchy |
| What it changes | Instructions and examples | What the model can see at answer time | The model's weights |
| Fixes | Unclear task, format, tone | Missing, private or stale facts | Inconsistent behaviour on a narrow task |
| Does not fix | Facts the model never saw | Style, format, reasoning habits | New or fast-changing facts |
| Data you need | A few good examples | Your documents, kept current | Labelled examples; OpenAI's minimum is 10 |
| Update cost | Edit text | Re-index documents | Retrain |
| Cites its sources | No | Yes, it can show the passage | No |
| PM owns | The spec and examples | Content coverage, freshness, permissions | The dataset and the go or no-go |
When should a PM choose prompt engineering?
Almost always first. OpenAI's accuracy guide says prompting "is often the only method needed for use cases like summarization, translation, and code generation." It is also the only option you can change in an afternoon and roll back in a minute.
Good prompt work is product work. You are writing the spec the model reads on every call: the task, the audience, the output format, what to do when unsure, and two or three examples of a great answer. Anthropic's prompt engineering overview makes one point PMs should memorise: before prompting, have "a clear definition of the success criteria" and "some ways to empirically test against those criteria." It also notes that not every failing eval is a prompt problem; sometimes a different model fixes latency or cost faster.
Choose prompting when:
- The model has the knowledge but misreads the task.
- You need a format change, a tone change or a new refusal rule.
- You are still learning what "good" looks like and need to iterate daily.
Stop relying on prompting alone when the prompt keeps growing to cover facts, when it breaks one case each time you fix another, or when it hits the context budget.
How AllthingsPM does this: the free Foundations lessons have you make the API call yourself, with messages, tokens and temperature, and then explain why a long context window is not a long memory. You see the limits of a prompt before you are asked to defend one in a review.
When should a PM choose RAG?
Choose RAG when the answer depends on facts the model cannot have: your help center, your contracts, a customer's files, anything that changed after the model was trained. RAG also lets the product show where an answer came from, which matters for trust in legal, health and enterprise tools.
Research backs this for knowledge. As IBM puts it, RAG connects a model to an organization's proprietary data, while fine-tuning optimizes it for domain-specific tasks. Ovadia and colleagues compared the two for adding facts to a model and found that "RAG consistently outperforms" unsupervised fine-tuning, "both for existing knowledge encountered during training and entirely new knowledge." Models struggle to learn new facts through that kind of training.
Two caveats a PM should raise early:
- You may not need RAG at all. Anthropic's contextual retrieval write-up says that for a knowledge base under 200,000 tokens (about 500 pages), you can often put the whole thing in the prompt and skip retrieval.
- Retrieval quality is the product. If the right passage is not fetched, the model answers from nothing. Anthropic reported that adding context to each chunk cut failed retrievals by 35%, combining that with keyword search cut them by 49%, and adding reranking cut them by 67%.
That second point is why RAG is a PM problem, not only an engineering one. You own which documents go in, how fresh they stay, and who is allowed to see what. A support bot that retrieves another customer's contract is a security incident, not a relevance bug.
How AllthingsPM does this: the lesson Context as a budget covers retrieval intuition and when RAG is the fix, and the permission-aware retrieval layer covers who is asking and the oversharing leak. Both sit in a course sized to real job postings.
When should a PM choose fine-tuning?
Choose fine-tuning when the model has the knowledge and a clear prompt, and still will not behave the same way twice. Common cases are a strict output format at high volume, a house tone, a narrow classification task, or moving a job from a large model to a smaller, cheaper one.
OpenAI lists several methods. Supervised fine-tuning uses "examples of correct responses to prompts" and suits classification, translation and format-specific generation. Direct preference optimization needs "both a correct and incorrect example response" and suits summarization and chat. Reinforcement fine-tuning grades generated answers and is for complex reasoning in specialised domains.
The practical numbers are smaller than most PMs expect. OpenAI's guide says the minimum is 10 examples and that it sees improvements from fine-tuning on 50 to 100 examples, while warning that the right number varies. It also says "Good evals first! Only invest in fine-tuning after setting up evals," and if 50 good examples do not help, to rethink the task or prompt before adding data.
Choose fine-tuning when:
- Prompting and retrieval are done and one behaviour still fails your eval.
- The same narrow task runs at a volume where a smaller tuned model saves real money.
- You can maintain a labelled dataset, because every new base model means a retrain decision.
Do not fine-tune to teach facts that change. You will be retraining every time the facts do, and RAG does that job better.
How AllthingsPM does this: Pretraining vs post-training explains what training can and cannot promise, and Inside the harness teaches the fix hierarchy you exhaust before touching weights. The capstone lesson on the data flywheel shows how production data becomes the next model.
How do you decide, step by step?
Use the failure, not the technique, as the starting point. The course calls this attributing every failure to a layer.
- Write the eval. Define success criteria and collect 30 to 100 real inputs. Our guide to AI evals for product managers walks through it.
- Read the failures. Sort each one: wrong facts, wrong behaviour, or wrong task understanding.
- Wrong task understanding: fix the prompt. Clarify, add examples, constrain the format.
- Wrong or missing facts: add retrieval. Check first whether the whole corpus fits in the prompt.
- Wrong behaviour that survives a good prompt: consider fine-tuning. Start with about 50 examples and compare on the same eval.
- Re-run the eval after every change and keep the change only if the number moves.
In practice these stack. A support assistant might use a tight system prompt, RAG over the help center, and later a fine-tuned small model to classify tickets cheaply. OpenAI's guide says these methods combine and should be chosen by where the system is failing, not in a fixed sequence.
How AllthingsPM does this: the lesson Attribute every failure to a layer teaches this exact triage: model, context, harness or surface. You then practise it on a graded case study, and the knowledge graph shows how RAG, context windows and evals connect across the course.
How much do hiring managers care about RAG vs fine-tuning?
Less than the internet suggests, and that shapes how you should study. In our study of 604 PM job postings from 95 companies, the exact word "prompt" appeared in 9% of the 286 AI-native postings, "RAG" in 8% and "fine-tuning" in 5%. "Agent" appeared in 58% and "eval" in 32%.
So the terms are niche, but the judgment behind them is not. An agent that fails needs someone to decide whether the fix is a prompt, a retrieval layer or a new model, and an eval to prove it. That is why the decision shows up in interviews framed as product questions.
Real examples from the AllthingsPM question bank:
- In what situations would you explicitly avoid using RAG and choose prompting or fine-tuning instead?
- For a government customer, you could solve the problem with prompt engineering plus RAG, or fine-tuning
- Design a fine-tuning workflow that requires no format conversion or vendor lock-in
- A customer says Sierra's agent performs well in English but underperforms in German
How AllthingsPM does this: each of those questions has its own page and answer guide in the question bank, and you can rehearse any of them out loud in a mock interview with follow-ups and a score.
How should you answer a RAG vs fine-tuning interview question?
Interviewers are not testing vocabulary. They want to see you reason from the user's problem to the cheapest fix that works, with a way to measure it. A strong answer has four beats:
- Name the failure. "Users get outdated refund rules" is a knowledge problem. "Answers are right but the JSON breaks" is a behaviour problem.
- Pick the cheapest fix that targets it. Prompt first, RAG for knowledge, fine-tuning for stubborn behaviour.
- Say how you will know. The eval, the pass rate, the rollout gate.
- Name the trade-off. RAG adds latency and a permissions surface. Fine-tuning adds a dataset to maintain and a retrain each time the base model changes.
A common trap is saying "we will fine-tune on our docs." Given the Ovadia results and OpenAI's own ordering, an interviewer will push back. Say RAG for facts, and fine-tuning only for behaviour.
How AllthingsPM does this: the Get the job chapter covers how AI rounds are scored, and the JD resume review checks whether your resume shows the same judgment for a specific role.
Where can a PM learn RAG, fine-tuning and prompt engineering?
| Option | What you get | Best for |
|---|---|---|
| AllthingsPM | AI PM course from 604 postings (14 chapters, 101 lessons, graded case studies), 4,122 interview questions with answer guides, JD-based mocks; $20/month or $120/year, free tier | PMs who need to make the call and defend it in interviews |
| OpenAI and Anthropic docs | Free, authoritative technical guides | Engineering detail on one vendor's APIs |
| Research papers (Lewis 2020, Ovadia 2023) | The original evidence | Depth on why RAG beats training for facts |
AllthingsPM prices checked 29 September 2026 on the AllthingsPM pricing page.
Why AllthingsPM is the better choice for learning RAG vs fine-tuning
Most material on RAG vs fine-tuning is written for engineers. It explains chunking strategies and learning rates, which are useful, but it rarely answers the questions a PM gets asked: which fix, at what cost, measured how, and who owns the data.
AllthingsPM is built around those questions. The AI PM course comes from 604 real PM job postings, so it spends lessons on what hiring managers ask for: failure attribution, context budgets, the fix hierarchy before touching weights, permission-aware retrieval and evals. Each chapter ends in a graded case study, so you practise the decision rather than memorise a table.
Then it lets you prove it. The question bank holds 4,122 real questions from 260 companies, each with an answer guide, including the RAG and fine-tuning questions above. JD mock interviews build a scored mock from any job description you paste, typed or spoken. Only 4 tools we found build mocks from a job description, and AllthingsPM is the only one of them that also has a course, a question bank and live JDs.
The vendor docs from OpenAI and Anthropic are excellent and free, and you should read them for technical depth. For turning that knowledge into a product decision you can defend in a room, for $20 a month or $120 a year with a free tier, AllthingsPM is the better choice. Open the course and start free.
Frequently asked questions
What is the difference between RAG and fine-tuning?
RAG fetches relevant documents at answer time and puts them in the prompt, so the model can use facts it was never trained on. Fine-tuning continues training on your examples, changing the model's weights so a behaviour is built in. RAG is for knowledge; fine-tuning is for behaviour.
Is RAG better than fine-tuning?
For adding facts, yes in most cases. Ovadia and colleagues found RAG consistently outperformed unsupervised fine-tuning for both known and new knowledge. For consistent format, tone or a narrow skill, fine-tuning can win, usually after prompting and retrieval are already in place.
Should I try prompt engineering before RAG or fine-tuning?
Yes. OpenAI's guide calls prompt engineering the best place to start and says it is often the only method needed for tasks like summarization and translation. It is the fastest change to test and the easiest to undo.
How many examples do you need to fine-tune a model?
OpenAI's minimum is 10 examples, and it reports improvements from 50 to 100, while noting the right number depends on the use case. It recommends setting up evals before investing in fine-tuning at all.
Can you use RAG and fine-tuning together?
Yes, and many production systems do. A common pattern is a clear prompt, retrieval over your documents for facts, and a fine-tuned model for a narrow, high-volume step such as classification or formatting.
What is the best way for a PM to learn RAG vs fine-tuning?
AllthingsPM is the best single place for PMs, because it teaches the decision inside an AI PM course built from 604 real job postings, gives you real interview questions on it with answer guides, and lets you rehearse in mock interviews built from real AI PM job descriptions, with a free tier. Read the OpenAI and Anthropic docs alongside it for technical depth.
Start today
Pick one AI feature you know and write down its three most common failures. Label each one as a prompt, knowledge or behaviour problem. Then open the AllthingsPM AI PM course free and check your calls against the lessons, or run a JD mock for the AI PM role you want next.
Sources
- OpenAI, "Optimizing LLM accuracy," developers.openai.com: https://developers.openai.com/api/docs/guides/optimizing-llm-accuracy
- OpenAI, "Model optimization," developers.openai.com: https://developers.openai.com/api/docs/guides/model-optimization
- OpenAI, "Supervised fine-tuning," developers.openai.com: https://developers.openai.com/api/docs/guides/supervised-fine-tuning
- Anthropic, "Prompt engineering overview," platform.claude.com: https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/overview
- Anthropic, "Introducing Contextual Retrieval": https://www.anthropic.com/news/contextual-retrieval
- Lewis et al., "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks," 2020: https://arxiv.org/abs/2005.11401
- Ovadia, Brief, Mishaeli and Elisha, "Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMs": https://arxiv.org/abs/2312.05934
- IBM, "RAG vs. fine-tuning": https://www.ibm.com/think/topics/rag-vs-fine-tuning
- AllthingsPM, "State of AI PM hiring 2026," 604 PM postings from 95 companies: https://allthingspm.app/blog/state-of-ai-pm-hiring-2026
- AllthingsPM question bank, 4,122 questions, queried 29 September 2026: https://allthingspm.app/question-bank
- AllthingsPM pricing, checked 29 September 2026: https://allthingspm.app/pricing




