An AI PRD is a product requirements document for a feature whose output is not fully predictable. It keeps the classic core (problem, users, scope, success) and adds six sections a normal PRD skips: an AI-or-not decision, a failure mode and risk register, layered guardrails, human escalation triggers, an eval plan that works as the acceptance criteria, and success metrics at three levels with cost and latency limits. The template below has all of them, and AllthingsPM is where you turn it into a skill: a five-lesson AI PRD chapter with a graded case, plus mock interviews built from any real AI PM job description. Copy the template, delete what your feature does not need, and never delete the risks.
AllthingsPM is an AI PM course and PM interview prep platform. The template is free, and everything in it is sourced.
What is an AI PRD, and how is it different from a normal PRD?
A normal PRD assumes deterministic software that QA checks once. An AI feature gives different answers to the same input and changes when a prompt or model changes, so the spec must describe a range of behaviour.
Miqdad Jaffer and Paweł Huryn published a widely shared AI PRD template in March 2025, noting that "AI-specific considerations are often overlooked." This version puts risks, guardrails and evals in the middle of the document, because on AI features they decide scope.
| Section | Classic PRD | AI PRD |
|---|---|---|
| Why build it | Problem and opportunity | Plus why AI beats a rule, a classifier or a form |
| Behaviour | Functional requirements | Capabilities, plus what the model must never do |
| Quality | Pass or fail acceptance criteria | Eval plan: golden set, graders, pass rate thresholds |
| Risk | A short risks line | Failure mode register with severity and owner |
| Safety | Usually none | Layered guardrails and human escalation triggers |
| Success | One or two KPIs | Outcome, task and step metrics, plus counter metrics |
| Cost | Engineering estimate | Cost per task, latency budget, spend caps |
Do employers actually expect AI PMs to write this?
Yes. In our study of 604 PM job postings from 95 companies, PM craft (PRDs, roadmaps, strategy) appeared in 88% of the 286 AI-native postings. What set them apart was the AI part of the spec: "eval" appeared in 32% of AI-native postings against 3% of other PM postings, and "safety" or "guardrails" in 12% against 4%.
Interviews test the same thing. Our question bank has real questions such as "After launching an enterprise AI agent, what primary success metric and guardrail metrics would you track?". That is the AI PRD, asked out loud.
How AllthingsPM does this: every question in the question bank has its own page and answer guide, so you can see what a strong guardrail and metrics answer looks like before you try one. Then open a live role in the jobs catalog and run a mock built from its exact description. You rehearse the spec for the job, not for a generic prompt.
The AI PRD template (copy it)
Paste this into your doc tool.
AI PRD: [feature name]
Owner: [PM] Eng lead: [ ] Design: [ ] Data/ML: [ ] Status: [draft]
Last updated: [date] Model(s): [name and version] Review: [legal, security, trust]
1. PROBLEM AND USERS
- Who has the problem, how often, what it costs them today (evidence)
2. WHY AI (AND WHY NOT SOMETHING SIMPLER)
- Alternatives considered: rules, search, a form, a classifier, a human
- Why they fall short
- Cost of a wrong answer for the user and for us (low / medium / high)
3. SCOPE
- In scope: tasks the model does
- Out of scope: tasks it must refuse or hand off
- Autonomy level: suggests / drafts for approval / acts on its own
4. USER EXPERIENCE
- Happy path, step by step
- What the user sees when the model is unsure, wrong, slow or down
- How the user corrects, undoes or reports an answer
5. FAILURE MODES AND RISK REGISTER
| Failure mode | Example | Severity | Likelihood | Guardrail | Eval | Owner |
6. GUARDRAILS
- Input: relevance, jailbreak and prompt injection, moderation, PII
- Tools and actions: risk rating per tool, permissions, spend or volume caps
- Output: format validation, PII, policy and brand checks, grounding
- Over-refusal budget: how often the guardrails may block a legitimate request
7. HUMAN IN THE LOOP
- Escalation triggers: failure thresholds, high-risk actions, low confidence
- Who receives the handoff, with what context, within what time
8. EVAL PLAN (ACCEPTANCE CRITERIA)
- Golden set: size, sources, who labels it
- Criteria: binary pass/fail per failure mode
- Graders: code checks, LLM judge (validated against human labels), human review
- Launch gate: minimum pass rate per criterion, zero tolerance items
9. SUCCESS METRICS
- Outcome (business): [metric, baseline, target, date]
- Task (per session): [task success rate, escalation rate, edit rate]
- Step (per model call): [eval pass rate, tool call accuracy]
- Counter metrics: what must not get worse
10. COST, LATENCY AND CAPACITY
- Cost per successful task, target and ceiling
- Latency budget (first token and full answer)
- Rate limits, spend caps, fallback model
11. DATA AND PRIVACY: what the model sees, retention, training use
12. ROLLOUT AND MONITORING: stages with exit criteria, online signals,
evals to rerun before any model or prompt change, kill switch owner
13. OPEN QUESTIONS AND DECISIONS LOG
How AllthingsPM does this: the template is the skeleton; the AI PRD chapter teaches each section with worked examples, and its graded case study asks you to write your own AI PRD and get feedback on it. That feedback loop is the part a copied template never gives you.
How do you write the "why AI" section?
This section exists to stop the PRD before it starts, when it should. Google's People + AI Guidebook says AI is "probably not better" for "Minimizing costly errors. If the cost of errors is very high and outweighs the benefits of a small increase in success rate". Anthropic's agent guidance agrees: "we recommend finding the simplest solution possible, and only increasing complexity when needed."
List the simpler options and why they fall short, rate the cost of a wrong answer, and state the autonomy level: drafting a reply needs far fewer guardrails than sending it. Our free lesson on the AI-or-not decision covers when SQL or a classifier beats an LLM, and the lesson on levels of autonomy covers where the human belongs.
A worked example: a team wants an LLM to route support tickets into ten fixed queues. The simpler option is a classifier trained on past tickets. It is cheaper per call, faster and easier to evaluate, so the "why AI" section should say so and either stop the PRD or narrow the LLM to the free-text tickets the classifier cannot handle. Writing that down is what interviewers mean by product judgment.
How do you write the risks section?
Write failure modes before features. One row per way the feature can hurt a user or the business, with an example, a severity, a guardrail, an eval and an owner. A row with no guardrail and no eval is not managed.
The OWASP Top 10 for LLM Applications 2025 is a good starting list, but the real rows come from reading your own outputs. Our AI evals guide shows how to read about 100 traces and group what went wrong; those groups become the register.
| Failure mode (example: support agent) | Severity | Guardrail | Eval |
|---|---|---|---|
| States a refund policy that does not exist | High | Answers grounded in the policy knowledge base, with a cited source | Grounding check on golden set |
| Issues a refund above the allowed limit | High | Refunds above limit need human approval | Code assertion on tool calls |
| Follows instructions hidden in a customer email | High | Safety classifier on inputs; tools cannot be triggered by message content alone | Red team set of injection attempts |
| Leaks another customer's personal data | High | PII filter on outputs; scoped data access | Code check for PII patterns |
| Refuses a legitimate billing question | Medium | Over-refusal budget; relevance classifier tuned on real queries | Pass rate on legitimate queries |
| Loops without resolving the issue | Medium | Retry cap, then hand off to a human | Task success and escalation rate |
These rows are illustrative.
How AllthingsPM does this: our AI evals guide and the course Evals chapter teach you to build this register from real traces, and the question bank lets you practise explaining it under interview pressure. A risk register you can defend out loud is worth more than one you copied.
How do you specify guardrails and human handoffs?
In layers. OpenAI's practical guide to building agents says: "Think of guardrails as a layered defense mechanism. While a single one is unlikely to provide sufficient protection, using multiple, specialized guardrails together creates more resilient agents." It lists seven types: relevance classifier, safety classifier, PII filter, moderation, tool safeguards, rules-based protections and output validation. Write one PRD line for each that applies.
For agents, rate every tool on "read-only vs. write access, reversibility, required account permissions, and financial impact," as the guide puts it. Then set an over-refusal budget, because every guardrail also blocks some legitimate users.
For handoffs, the same guide names two triggers: exceeding failure thresholds, and high-risk actions such as "canceling user orders, authorizing large refunds, or making payments." Write both as numbers (the retry count before handoff, the list of actions that always need approval), then say who receives the handoff and with what context. The course lessons on guardrails in the request path and guardrails in the PRD cover both sides.
How AllthingsPM does this: guardrail design is one of the most common AI PM interview prompts, so we practise it three ways: the two course lessons above, real guardrail questions in the question bank with answer guides, and a JD mock where the AI interviewer pushes back with follow-ups like "what does this guardrail block by mistake?"
How do you write the eval plan?
Treat it as the acceptance criteria. Anthropic recommends "extensive testing in sandboxed environments, along with the appropriate guardrails" before agents act for real. A good plan names four things:
- The golden set: size, sources (real traces first, synthetic to fill gaps), and who labels it.
- The criteria: binary pass or fail per failure mode in the register.
- The graders: code checks where a rule can see the failure, an LLM judge validated against human labels where it cannot, humans for the rest.
- The launch gate: minimum pass rate per criterion, plus zero tolerance items such as no PII leaks on the red team set.
The Evals chapter of the course goes deeper, from reading traces to building golden datasets.

What success metrics and cost limits belong in an AI PRD?
Three levels, because a feature can pass its evals and still fail the business:
- Outcome: did the business result move? Resolution without a human, time saved, conversion.
- Task: did the feature finish the user's job? Task success, escalation, edit or regenerate rate.
- Step: was each model call right? Eval pass rate, tool call accuracy, grounding rate.
Give each a baseline, target and date, and add counter metrics that must not get worse, such as complaint rate or refund reversals.
Cost and latency are product decisions too. Anthropic notes that "Agentic systems often trade latency and cost for better task performance." Put three numbers in the PRD: target cost per successful task (failed attempts cost money too), the latency budget, and a spend cap with a fallback. The lesson on cost per successful task and latency shows how to set them.
A worked example for the support agent above: outcome metric, share of tickets resolved without a human; task metric, escalation rate and customer edit rate on drafted replies; step metric, grounding pass rate on the golden set; counter metric, refund reversals. Each one gets a baseline, a target and a date.
How AllthingsPM does this: metrics questions like "what success and guardrail metrics would you track?" come up again and again in AI PM interviews. Our answer guides show how to structure the three levels, and a JD mock scores whether you actually did it when you answer out loud.
A mini plan you can run today on AllthingsPM
- Open the AI PRD chapter and read the first lesson (20 minutes).
- Fill in sections 2, 5 and 9 of the template for a feature you know well.
- Answer one guardrail question and one metrics question from the question bank, then compare with the answer guides.
- Pick a role in the jobs catalog and run a mock built from its description, in voice if you can.
- Fix the weakest section of your PRD based on the feedback, and repeat tomorrow.
Why AllthingsPM is the better choice for learning the AI PRD
Your goal is not a document; it is being able to write one at work and defend it under follow-up questions in an interview. A free template (including the good one from Product Compass) gives you the skeleton. A generic AI chat tool will draft sections for you, but it will not grade your judgment against a real role. AllthingsPM is the only tool we found that puts the lesson, the graded practice and the real roles in one place instead of five:
- The course is built from 604 real PM job postings, with 14 graded case studies, including one on your own AI PRD.
- The question bank has 4,122 real questions from 260 companies, and every question has an answer guide.
- The jobs catalog lists live AI PM roles, such as the Abridge evals role, each with a mock built from the exact job description.
| Option | Teaches the AI PRD | Graded practice | Mock from a real JD | Price |
|---|---|---|---|---|
| AllthingsPM | Yes, 5 lessons plus a graded case | Yes | Yes, text or voice | Free tier; $20 a month or $120 a year |
| Product Compass template (Mar 2025) | Template with guidance | No | No | Free article |
| Generic AI chat tool | Drafts on request | No structured grading | No | Varies |
| Books and blogs | Concepts | No | No | Varies |
The fair line on each: the Product Compass template is a strong, free skeleton; a chat tool is fast for drafting; books build depth. None of them grades your AI PRD or tests it against a real job description. For learning to write and defend an AI PRD, AllthingsPM is the clear choice. Open the AI PRD chapter.
Frequently asked questions
What is an AI PRD?
An AI PRD is a product requirements document for a feature built on an AI model. It keeps the problem, users, scope and success sections of a normal PRD and adds an AI-or-not decision, a failure mode register, layered guardrails, human escalation triggers, an eval plan used as acceptance criteria, and cost and latency limits.
How is an AI PRD different from a traditional PRD?
A traditional PRD specifies one expected output per input and checks it once. An AI PRD specifies a range of acceptable behaviour, because the model gives different answers to the same input. Its acceptance criteria are pass rates on a golden set, and risks and guardrails sit in the middle of the document.
What guardrails should an AI PRD include?
Use layers: input checks (relevance, jailbreak and prompt injection, moderation), tool safeguards with a risk rating per tool, and output checks (PII, format, policy). OpenAI's agent guide lists seven types and says a single guardrail is unlikely to be enough. Add an over-refusal budget.
What success metrics should an AI feature have?
Define them at three levels: an outcome metric for the business, task metrics such as task success and escalation rate, and step metrics such as eval pass rate and tool call accuracy. Give each a baseline, a target and a date, and pair it with counter metrics.
What is the best way to learn to write an AI PRD?
AllthingsPM first: its AI PRD chapter teaches every section and grades your own AI PRD in a case study, the question bank has real guardrail and metrics questions with answer guides, and JD mocks test you against real roles. Free templates such as Product Compass are a good skeleton to read alongside.
How do I learn to write an AI PRD for interviews?
Copy the template above, then take the AI PRD chapter on AllthingsPM, which grades your own AI PRD in a case study. Practise saying it out loud with a JD mock interview built from a real AI PM role.
Start today, free
You now have the template. The fastest way to make it yours is to write one, get it graded and defend it out loud. Start free on AllthingsPM: open the AI PRD chapter, answer one guardrail question, and run your first JD mock before the day is out.
Sources
- Miqdad Jaffer and Paweł Huryn, "A Proven AI PRD Template," Product Compass, 26 March 2025. https://www.productcompass.pm/p/ai-prd-template
- OpenAI, "A practical guide to building agents." https://cdn.openai.com/business-guides-and-resources/a-practical-guide-to-building-agents.pdf
- Anthropic, "Building effective agents," 19 December 2024. https://www.anthropic.com/engineering/building-effective-agents
- Google PAIR, People + AI Guidebook, "User Needs + Defining Success." https://pair.withgoogle.com/chapter/user-needs/
- OWASP, "Top 10 for LLM Applications 2025." https://genai.owasp.org/llm-top-10/
- AllthingsPM, "State of AI PM Hiring 2026: What 604 Job Postings From 95 Companies Ask For." https://www.allthingspm.app/blog/state-of-ai-pm-hiring-2026




