Flash sale 30% off with code LAUNCH30 Ends in --:--:--
See pricing
All Things PM

AI Safety and Guardrails for PMs: A Practical Guide (2026)

AI guardrails are the checks around a model that block bad inputs, bad outputs and bad actions. PMs own which ones ship, what they cost in refusals, and who approves irreversible actions. AllthingsPM teaches this in chapter 12 of its AI PM course.

AllthingsPM·September 29, 2026·16 min read
A product manager at a desk with a highlighted printed job description next to a laptop showing a half-built app prototype
Guardrails do not slow the product down. They let it drive faster near the edge.

AI guardrails are the checks that sit around a model and stop bad inputs, bad outputs and bad actions before they reach a user or the real world. For a product manager, guardrails are a product decision, not a security afterthought: you decide which risks the product must never take, how many good requests you are willing to refuse to avoid them, and which actions need a human to approve. The fastest way to learn this as a working skill is the Trust, safety, and agent security chapter of the AllthingsPM AI PM course, which turns each guardrail into a PRD line, a metric and an interview answer.

AllthingsPM is an AI PM course and PM interview prep platform. Its course is built from 604 real PM job postings, and this guide follows the same order as its lessons.

What are AI guardrails, in product terms?

A guardrail is any rule, classifier, filter or approval step that runs outside the model's own judgment and can change what happens next. The model may still try to do the wrong thing. The guardrail makes sure it does not land.

Engineering frameworks describe the same idea in layers. NVIDIA's open source NeMo Guardrails, for example, separates input rails, retrieval rails, dialog rails, execution rails and output rails, each running at a different stage of one interaction [5]. OpenAI's Agents SDK has input guardrails, output guardrails and tool guardrails, and any of them can trip a "tripwire" that halts the run immediately [3].

For a PM, the useful version is five questions, one per layer:

LayerThe PM questionExample guardrailFailure it stops
InputWhat should we refuse to even process?Jailbreak and abuse classifier on the user messagePrompt injection, policy abuse
RetrievalWhat content can the model trust?Only index approved sources; tag web content as untrustedPoisoned or stale context
Action (tools)What can the model do without asking?Allow-list of tools, spend limits, human approval for irreversible stepsExcessive agency, data exfiltration
OutputWhat must never reach the user?PII scrubber, policy check, grounding check against sourcesLeaks, harmful or false answers
RuntimeHow do we notice when it breaks anyway?Logging, alerts, rate limits, kill switch, incident runbookSilent drift, runaway cost

The table is the core of the job. Every AI feature you spec should have at least one named guardrail in each row, or a written reason why that row does not apply.

How AllthingsPM does this: the lesson Guardrails in the request path, and the over-refusal budget on the other side walks through these layers on a real feature, then asks you to write the guardrail spec yourself. It sits inside the AI PM course, so the same agent you design in earlier chapters is the one you secure here.

Why do product managers own guardrails and not just engineers?

Because the hardest guardrail questions are tradeoffs, and tradeoffs are product calls. Engineers can build a classifier that blocks 99 percent of abuse. Only the PM can say whether blocking it is worth refusing a slice of honest users.

Liability also lands on the company, not the model. In February 2024, British Columbia's Civil Resolution Tribunal held Air Canada responsible after its website chatbot told a customer he could claim a bereavement fare retroactively, which the airline's policy did not allow. The tribunal rejected the idea that the chatbot was responsible for its own words and awarded the customer $650.88 in damages [4]. The amount was small. The lesson for PMs was not: whatever your AI says, your company said.

Hiring reflects this. In the AllthingsPM corpus of 389 AI PM postings at 86 companies, read in September 2026, 63 postings mention safety and 15 mention guardrails by name. Anthropic alone lists roles such as Product Manager, Multi-Cloud Trust & Safety, and OpenAI lists Product Manager, Safety Measurement and Product Manager, Multimodal Safety.

Bar chart led by AllthingsPM (us), which teaches all six topics in its AI PM course, then the number of 389 AI PM job postings that mention each term: safety 63, evals 42, red teaming 17, guardrails 15, responsible AI 4, prompt injection 2
AllthingsPM JD corpus, 389 AI PM postings at 86 companies, keyword match, September 2026

Read the chart this way: safety and evals are already mainstream PM requirements, and the more specific terms (red teaming, prompt injection) show up mostly in roles at labs and agent companies, where interviewers go deepest.

How AllthingsPM does this: every role in the jobs catalog opens a mock interview built from that exact job description, so you can rehearse a safety PM loop at Anthropic or OpenAI before you apply. The JD mock does the same for any posting you paste in.

What risks should your guardrails cover?

Start from a shared list instead of a blank page. The OWASP Top 10 for LLM Applications (2025 edition) is the most widely used one [1]. The ten items are prompt injection, sensitive information disclosure, supply chain, data and model poisoning, improper output handling, excessive agency, system prompt leakage, vector and embedding weaknesses, misinformation, and unbounded consumption.

You do not need all ten in every PRD. Map them to your feature:

  • A support chatbot mostly worries about misinformation, sensitive information disclosure and system prompt leakage.
  • A RAG search product adds vector and embedding weaknesses and data poisoning.
  • An agent that takes actions adds excessive agency, improper output handling (its output becomes someone else's input) and unbounded consumption, because loops cost money.

For governance, NIST's AI Risk Management Framework gives four functions, Govern, Map, Measure and Manage, and a Generative AI Profile (NIST AI 600-1, released July 2024) that lists risks specific to generative systems [2]. Treat it as a checklist for what your company needs, not a document to reproduce.

How AllthingsPM does this: the lesson Right-sized governance is explicit about which frameworks matter for a PM and which you can skip, so you spend time on the guardrails that change outcomes instead of paperwork.

How do you balance safety against over-refusal?

Every guardrail has two error types. It can let a harmful request through, or it can block a perfectly good one. The second error is the one teams forget to measure, and it is the one users feel every day.

A practical way to manage it is an over-refusal budget, written into the PRD next to the safety target:

  1. Pick the harm metric. For example, the share of red-team prompts in a golden set that produce a policy violation.
  2. Pick the refusal metric. The share of legitimate prompts in a separate benign set that get refused or watered down.
  3. Set both targets together. "Violation rate under X on the adversarial set, false refusal under Y on the benign set." A guardrail change that improves one and breaks the other does not ship.
  4. Add latency and cost. Blocking guardrails run before the model; parallel ones run alongside it. OpenAI's SDK documents exactly this choice: blocking mode prevents wasted tokens, parallel mode gives the best latency [3].

This is where guardrails meet evals. The same golden sets and graders you use for quality measure your safety layer.

How AllthingsPM does this: the course's Evals chapter teaches you to build golden sets and defensible numbers, and the Trust chapter reuses them for safety, so you learn one measurement system instead of two. For a deeper read, see AI evals for product managers.

What changes when your product is an agent?

Chatbots say things. Agents do things. That moves the most important guardrails from the output layer to the action layer.

Simon Willison's "lethal trifecta," published in June 2025, is the clearest mental model a PM can carry [6]. An agent is at serious risk when it combines three things:

  1. Access to private data, such as email, files or a CRM.
  2. Exposure to untrusted content, such as web pages, inbound email or user uploaded documents.
  3. The ability to communicate externally, such as sending email, calling an API or rendering a link.

In his words: "If your agent combines these three features, an attacker can easily trick it into accessing your private data and sending it to that attacker" [6]. Prompt injection is how the trick works: instructions hidden in untrusted content get treated as if the user wrote them.

The PM move is to break the trifecta on purpose. Remove one leg for the risky flows: read-only mode when untrusted content is in context, no outbound calls while private data is loaded, or a human approval step before anything leaves the system.

A short agent guardrail checklist for your spec:

  • Tool allow-list with a risk level per tool (read, write, irreversible).
  • Human approval for every irreversible or money-moving action.
  • Scoped permissions so the agent has only the access this task needs.
  • Spend and step limits so a loop cannot run forever.
  • Audit log of every tool call, readable by support and security.

How AllthingsPM does this: the lesson Agent security: the lethal trifecta, prompt injection, tool misuse, and permission escalation applies this model to the agent you build across the course, and The agent spec you own covers scope, tool contracts, risk levels and escalation.

How do you write guardrails into a PRD?

If the guardrail is not in the PRD, it is optional, and optional guardrails get cut the week before launch. Put a dedicated section in the spec with four parts:

  • Never events. The short list of outcomes that block launch if they happen at all (for example, sharing another user's data, or completing a payment without confirmation).
  • Layered controls. One line per layer from the table above, naming the control and its owner.
  • Targets. The harm rate and over-refusal budget, measured on named eval sets.
  • Escalation and approval. Which actions need a human, who that human is, and how fast they must respond.

Write it before the demo impresses everyone. Once a prototype works, pressure to ship grows and safety turns into a follow-up ticket.

How AllthingsPM does this: two lessons cover exactly this: The AI PRD: name the risks, guardrails, and success metrics before you build and Guardrails in the PRD, and the human approval that makes an irreversible action safe. The companion post AI PRD template with guardrails gives you a copyable section.

How do red teaming and incident response fit in?

Guardrails are a hypothesis. Red teaming tests it. Before launch, a small team (or a vendor, or your own staff for a day) tries to break the product along the risks you listed: jailbreaks, injection through documents, requests for other users' data, cost blowups.

Keep it practical:

  1. Red-team the obvious modes first. Most real incidents come from simple attacks, not exotic ones.
  2. Turn every successful attack into an eval case. The golden set grows, and the same bug cannot quietly return.
  3. Write the incident runbook before launch. Who can switch the feature off, how you notify users, and how you tell a model regression from an attack.
  4. Watch for slow failure. Not every incident is an attack. Sometimes answers just get worse after a model update, and only monitoring catches it.

How AllthingsPM does this: the lesson Red-team the obvious modes, and the incident runbook for when the answers just get worse covers both halves, and the chapter ends with a graded integration case where you write the reliability number and blast radius memo for your agent.

How do guardrails come up in PM interviews?

At AI companies, often and directly. Interviewers ask you to design guardrails for a named product, to weigh a growth feature against a safety risk, or to hold a position on AI risk in a values round. Examples from the AllthingsPM question bank:

A strong answer follows the structure of this guide: name the users and the never events, walk the five layers, state the over-refusal budget, break the trifecta for agent actions, and close with how you would measure and red-team it.

How AllthingsPM does this: every question above has its own page with an answer guide and a one-click mock in the question bank, and the Anthropic company hub collects the lab's questions in one place. The course lesson on the values round prepares you to hold a position on AI risk under follow-up questions. If you are aiming at a lab, read safety and safeguards PMs at AI labs next.

Why AllthingsPM is the better choice for learning AI guardrails

Most material on guardrails is written for engineers. OWASP, NIST, NeMo Guardrails and the OpenAI Agents SDK docs are excellent references, and every source in this guide is worth bookmarking [1][2][3][5]. But none of them tells a PM how to set an over-refusal budget, write never events into a PRD, or answer "design guardrails for Operator" in a 45-minute interview.

AllthingsPM connects all three jobs. The AI PM course teaches guardrails as a product skill inside a curriculum built from 604 real PM job postings, with a full chapter on trust, safety and agent security and a graded case study at the end. The question bank gives you real guardrail questions from companies like Anthropic and OpenAI, each with an answer guide and a scored mock. The jobs catalog holds live safety PM roles, each with a mock built from the job description. General PM courses and bootcamps can teach frameworks and give you a cohort; for AI safety specifically, one account that teaches, drills and tests the skill against real roles is the more direct path.

The verdict: read the engineering references for depth, and use AllthingsPM to turn guardrails into something you can spec, measure and defend in an interview. Open the Trust chapter and start free.

Frequently asked questions

What are AI guardrails?

AI guardrails are checks that run outside the model and can block or change what happens: input filters, retrieval rules, tool permissions, output checks and runtime monitoring. They exist because the model itself can be tricked or wrong. Good products layer several of them.

What is the best way for a PM to learn AI guardrails?

AllthingsPM is the best starting point for PMs: its AI PM course has a full chapter on trust, safety and agent security, linked to real guardrail interview questions and live safety PM roles. Pair it with the OWASP Top 10 for LLM Applications and the NIST AI RMF as references.

Are guardrails the same as AI safety?

No. AI safety is the broader goal of making systems behave well, including model training and alignment work. Guardrails are the product-level controls a team adds around a model it ships. PMs usually own the guardrails, not the model's training.

What is over-refusal?

Over-refusal is when a guardrail blocks or waters down a legitimate request. It hurts user trust and retention as much as a harmful answer hurts safety. Measure it on a benign eval set and give it a budget in the PRD.

What is prompt injection?

Prompt injection is when instructions hidden in content the model reads, such as a web page or document, get treated as commands. OWASP lists it first in its 2025 Top 10 for LLM applications [1]. For agents, the fix is mostly architectural: limit what untrusted content can trigger.

Do agents need different guardrails than chatbots?

Yes. Agents take actions, so the most important controls move to the tool layer: allow-lists, scoped permissions, spend limits and human approval for irreversible steps. Breaking the lethal trifecta for risky flows is the core design move [6].

Which frameworks should a PM know?

Know the OWASP Top 10 for LLM applications for risks, the NIST AI RMF and its Generative AI Profile for governance, and one implementation framework such as NeMo Guardrails or the OpenAI Agents SDK so you can talk to engineers. You do not need to master all of them.

Ready to make guardrails a skill you can show? Start the AllthingsPM AI PM course free, then practice a guardrail design question as a scored mock.

Sources

  1. OWASP GenAI Security Project, "OWASP Top 10 for LLM Applications 2025": https://genai.owasp.org/llm-top-10/
  2. NIST, "AI Risk Management Framework" (AI RMF 1.0, January 2023; Generative AI Profile NIST AI 600-1, July 2024): https://www.nist.gov/itl/ai-risk-management-framework
  3. OpenAI Agents SDK documentation, "Guardrails": https://openai.github.io/openai-agents-python/guardrails/
  4. American Bar Association, "BC Tribunal Confirms Companies Remain Liable for Information Provided by AI Chatbot" (Moffatt v. Air Canada, February 2024): https://www.americanbar.org/groups/business_law/resources/business-law-today/2024-february/bc-tribunal-confirms-companies-remain-liable-information-provided-ai-chatbot/
  5. NVIDIA, NeMo Guardrails documentation: https://docs.nvidia.com/nemo/guardrails/latest/index.html
  6. Simon Willison, "The lethal trifecta for AI agents: private data, untrusted content, and external communication," 16 June 2025: https://simonwillison.net/2025/Jun/16/the-lethal-trifecta/
  7. AllthingsPM JD corpus, 389 AI PM postings at 86 companies, read September 2026 (keyword counts); AllthingsPM course, question bank and jobs catalog: https://allthingspm.app/course
PM
Written by the AllthingsPM team
Frameworks and interview prep for product managers.
The AI PM course

Reading is the easy half.
The course grades the other half.

Start for free