Flash sale 30% off with code LAUNCH30 Ends in --:--:--
See pricing
All Things PM

What Does an AI Product Manager Do?

An AI product manager owns a product whose core behavior comes from a model: they define good output and measure it with evals, trade off model quality, cost and latency, spec agents, design for wrong answers and get it deployed. AllthingsPM teaches and tests exactly this work, built from 604 real PM job postings.

AllthingsPM·September 26, 2026·16 min read
A product manager at a desk reading a long printed transcript with a pencil, marking some lines, while a laptop and a small stack of test cards sit beside them
Much of the job is reading what the model actually did, then deciding what good looks like.

An AI product manager owns a product whose core behavior comes from a model. They still do the classic PM work (customers, strategy, roadmap, launches), but five things are new: they define what good output means and measure it with evals, they make model trade-offs between quality, cost and latency, they decide what an agent may do on its own and where a human steps in, they design the UX for a system that is wrong some of the time, and, in most roles, they get the product deployed inside a customer's company. In the 286 AI-native PM postings we read in full, technical fluency appears in 90 percent, agents in 73 percent and evals in 53 percent.

AllthingsPM is an AI PM course and PM interview prep platform, and its course is built from the same 604 job postings this article draws on. That is the concrete advantage: instead of reading about the job, you practice each responsibility below against real roles. Every responsibility is a chapter in our AI PM course, every posting quoted is a live job description with its own mock interview, and you can start free today.

What does an AI product manager do, in one paragraph?

An AI PM decides what an AI product should do, for whom, and how you will know it works. The "how you will know" part is where the job differs. A normal feature either works or has a bug. A model-powered feature is right some percentage of the time, changes when the model changes, and costs money on every call. So the AI PM turns fuzzy goals ("the agent should handle refunds well") into a written quality bar, a set of test cases, a cost budget and a plan for when the model is wrong.

OpenAI's Core Models posting says it directly: the PM will "define success across offline evaluations and online product metrics, balancing model quality, usefulness, latency, safety, reliability, and cost." Read the full Core Models job description and run the mock built from it.

How AllthingsPM does this. The first lesson of our course, the AI PM job now, is free and maps the role to the postings. Then the Core Models page above lets you rehearse the exact interview that role implies, with an AI interviewer that asks follow-ups and scores you. Reading the description takes five minutes; answering questions about it out loud is what gets you hired.

What are the core responsibilities of an AI PM?

Seven responsibilities recur across the postings. The percentages are the share of 286 AI-native postings that ask for the theme, from our study of 604 PM job postings. The second column shows where you practice each one on AllthingsPM.

ResponsibilityWhere to practice it on AllthingsPMWhat it looks likeFrom a real postingShare of AI-native postings
Product craftThe AI PRDDiscovery, strategy, roadmap, spec, launchGlean: "Developing key parts of our product roadmap, marrying customers' needs with our product vision"88%
Model trade-offsCost per successful taskChoose models, weigh quality against cost and latencyGlean AI Quality: "own projections of LLM usage, cost, and capacity planning"90% (technical fluency), 20% (model cost named)
Evals and qualityEvals chapterDefine "good", build the test set, set launch gatesAbridge: "make evaluating a new model cheap and routine as frontier models ship frequently"53%
AgentsAgents and agentic architectureWorkflow or agent, tools, permissions, escalationOpenAI API Agents: "make building agents faster, more intuitive, more reliable, and more powerful"73%
AI UX and oversightAI UX and human oversightConfidence, citations, approval steps, a manual pathSierra: agents "that handle thousands of customer conversations a day"78%
Enterprise deploymentShipping into somebody else's companySecurity review, integrations, contract to productionScale AI: "Drive deployments from contract to production"79%
Working with researchWorking with researchersTurn failures into research prioritiesAnthropic Labs: "Work with researchers to understand what's emerging and what it means for users"Concentrated at labs

Defining "good" and measuring it (evals)

This is the most distinctive part of the job. An eval is a repeatable test of the model's output: realistic inputs, a written criterion for a good answer, and a way to grade each one (a rule, a person, or another model as judge). The PM usually owns the criterion, because it encodes what users need.

Hamel Husain, who teaches a widely taken evals course with Shreya Shankar, recommends "reviewing at least 100 traces" and suggests the owner of that judgment be "a domain expert or PM who understands user needs." Kevin Weil, then OpenAI's chief product officer, gave the reason on Lenny's Podcast: "The AI model you're using today is the worst AI model you'll ever use." Without an eval suite you cannot tell whether the next model is better for your users.

A worked example. Say you own a support agent that answers billing questions. A useful first eval might be 50 real (anonymized) customer messages, each with a one-line criterion such as "states the correct refund window and links the policy" and a pass or fail grade. You agree before the build that the agent ships when it passes, say, 45 of 50, and you rerun the same 50 every time the prompt or the model changes. That single page is worth more in an interview than any definition of "precision".

How AllthingsPM does this. Our evals chapter walks through writing criteria, building a test set and setting a launch gate, with graded case studies. Then practice the interview version: the Abridge evals posting and the Glean AI Quality posting each have a mock built from their text.

Making model trade-offs

A bigger model may answer better but cost more per call and respond more slowly; a smaller one may be fine for most requests. The PM decides which requests go where and what "fast enough" and "cheap enough" mean, because those numbers set the gross margin. The useful unit is cost per successful task, not cost per call: a cheap model that fails half the time and triggers a retry or a human handoff can end up the expensive choice.

How AllthingsPM does this. The cost per successful task lesson teaches routing decisions with numbers, and the question bank has trade-off questions reported at AI companies, so you rehearse the answer, not just the idea.

Speccing agents and their limits

OpenAI's practical guide defines agents as "systems that independently accomplish tasks on your behalf." The PM writes down what the agent may do, which tools it can call, and when it must hand over to a person. OpenAI's guide flags "high-risk actions" such as "authorizing large refunds." A good agent spec reads like a job description for software: goal, allowed tools, spending or action limits, what it must never do, and the exact trigger for escalation.

How AllthingsPM does this. Our lesson on workflow or agent covers when a fixed workflow beats an agent. Then run the mock for OpenAI, API Agents or Decagon, Senior Agent PM to test yourself on the questions those teams care about.

Designing for a product that is sometimes wrong

Microsoft's Guidelines for Human-AI Interaction capture the checklist: "Make clear how well the system can do what it can do," "Support efficient correction," and "Scope services when in doubt." In practice the AI PM decides where citations appear, when to show uncertainty, where the approval step sits and what the manual fallback is.

How AllthingsPM does this. The AI UX and human oversight chapter turns those guidelines into design decisions you can defend in a product sense round.

Getting it deployed inside a customer's company

79 percent of AI-native postings ask for enterprise and deployment work, and 21 use the phrase "forward deployed", which appeared in no non-AI posting. Scale AI's posting is explicit: "This is not a roadmap PM, a CSM, or a solutions engineer." The work is security reviews, integrations with the customer's systems, and proving value inside their workflow before the contract renews.

How AllthingsPM does this. Shipping into somebody else's company covers the contract to production path, and the Scale AI forward deployed posting has a mock built from it.

How is an AI PM different from a regular product manager?

Moment in the weekAI PMRegular PM
Writing the specFeatures plus the quality bar, test cases, failure modes, cost budget and escalation rulesFeatures, flows and acceptance criteria
Deciding it is readyThe eval pass rate clears a bar agreed before the buildQA passes, no blocking bugs
Triaging a problemFind a pattern across many transcripts; rerun the suite after a fixReproduce the bug, fix it, it stays fixed
A vendor ships an updateA new model can change quality, cost and speed overnight; test firstUsually ignore it
Designing the UXShow confidence and sources, add approvals, keep a manual pathThe software does what the button says

How AllthingsPM does this. If you are moving from a regular PM role, you already have the right column. The course is built to close the left column, and our AI PM vs PM guide shows how to tell that story in interviews.

What kinds of AI PM roles exist?

"AI PM" covers several quite different jobs. We grouped the 286 AI-native postings by title and product description.

Bar chart, with AllthingsPM (us) highlighted first as the place to mock each of these postings, of 286 AI-native PM postings grouped by what the PM owns: AI products and agents for end users 82 (29%), developer platform and infrastructure 81 (28%), models, evals and research 41 (14%), safety, trust and security 33 (12%), forward deployed and agent deployment 31 (11%), growth and monetization 18 (6%)
Source: AllthingsPM JD corpus, 286 AI-native PM postings from 95 companies, read 6 September 2026. Grouped by keywords in title and product description

Pick the shape that fits your background before you pick a company, for example OpenAI, API Agents (developer platform) or Sierra, Agent Development (forward deployed). Each posting has a mock built from its text, and our company pages collect reported questions for OpenAI, Anthropic, Sierra and Scale AI, part of 4,122 questions from 260 companies, each with an answer guide.

How AllthingsPM does this. The jobs catalog lets you browse live roles by company, so you can see which shape each team is really hiring for before you spend a week preparing. When you find a fit, run a Resume Job Match or a JD resume review against that exact posting.

How do you get ready for this job? A plan you can run today

A cheap test: build a small agent for a task you know well, write 30 test cases, and grade the outputs yourself. If that afternoon felt like the interesting part, the job suits you. The PM as builder chapter walks through exactly that. Here is a one-week plan inside AllthingsPM:

  1. Day 1: read the free lesson the AI PM job now and pick one role shape from the chart above.
  2. Day 2: open two live postings of that shape in the jobs catalog and list the responsibilities they share.
  3. Day 3: work through the evals chapter and write a 30-case eval for a product you use.
  4. Day 4: answer three AI questions from the company pages in the question bank out loud, then compare against each answer guide.
  5. Day 5: run a JD mock interview for your top posting, by voice if you can.
  6. Day 6: run a JD resume review against the same posting and rewrite two bullets around evals or agents.
  7. Day 7: rerun the mock and compare scores.

Why AllthingsPM is the better choice for learning what an AI PM does

Your goal is not to know the job description; it is to do this work well enough to be hired for it. That takes practice on the responsibilities above, graded, against real roles. Most readers stitch that together from blogs, AI chat tools, prep sites and coaches. AllthingsPM is the only tool we found that puts the whole loop in one place: a course built from real postings, mocks built from any JD, a question bank with answer guides, and live roles at AI companies.

What you get for this jobBuilt from real AI PM postingsMock for a specific roleCost
AllthingsPM14-chapter, 101-lesson course, JD mocks (text or voice, scored), 4,122 questions with answer guides, JD resume reviewYes, 604 postings, updated weeklyYes, any JD, plus 116 live AI company JDsFree tier; $20/month or $120/year
Blogs and booksConcepts and frameworksRarelyNoFree to low
General AI chat toolsPractice on whatever you paste inNoOnly if you write the promptVaries
Human coachesExpert, personal feedbackDepends on coachUsually one session at a timePer session

Each alternative has a real strength. Blogs and books, including our own 111 book summaries, are great for concepts. General chat tools are flexible. A coach gives the most personal feedback. But none of them gives you a curriculum built from the postings, a scored mock for the exact role and the questions those companies ask, all in one place. Verdict: for daily practice on evals, agents and trade-offs, use AllthingsPM, and save a coach for the final round if you want one. Start by pasting a posting into the JD mock interview.

Frequently asked questions

What is the main job of an AI product manager?

To decide what an AI product should do and prove that it does it well enough: the usual PM work plus a quality bar measured with evals, model choices on quality, cost and speed, and deciding where a human stays in the loop.

What is the best way to learn what an AI product manager does?

AllthingsPM, because it teaches and tests the job from the job itself: a 14-chapter AI PM course built from 604 real postings, mocks built from any job description, and 4,122 questions from 260 companies with answer guides, with a free tier and Pro at $20/month. Reading postings and blogs helps, but graded practice against real roles is what transfers to interviews.

Do AI product managers need to code?

They do not write production code, but most need to read code and build rough prototypes. Claude Code, Cursor, Codex or Copilot appear in 9 percent of AI-native postings against 1 percent of others.

Is an AI PM the same as a machine learning PM?

Not quite. ML PMs traditionally owned models their company trained, such as ranking or fraud models. Most AI PMs today build on foundation models, so the work shifts to evals, model selection, agent design and deployment.

What skills do AI product manager job descriptions ask for most?

Technical fluency (90 percent), PM craft (88 percent), enterprise deployment (79 percent), AI UX (78 percent), agents (73 percent) and evals (53 percent). Our AI PM interview questions guide shows how interviewers test each one.

How do I become an AI product manager?

Learn model basics and evals, build and measure one small AI product, and aim for the role shape your background makes credible. Our complete roadmap has the full sequence, and our comparison of AI PM courses shows why we rank the AllthingsPM course first.

Start today

Ready to do the job, not just read about it? Start free on AllthingsPM: open the first lesson, pick one real AI PM posting from the jobs catalog, and run its mock interview today. In about 30 minutes you will know exactly which of the responsibilities above you can already defend and which to practice next.

Sources

  1. AllthingsPM JD corpus: 604 PM postings from 95 companies read in full from Greenhouse, Ashby and Lever careers boards on 6 September 2026, of which 286 were AI-native. Method in State of AI PM Hiring 2026. Role-shape grouping computed for this article.
  2. AllthingsPM jobs catalog, job descriptions quoted above, checked 26 September 2026: OpenAI, Core Models; OpenAI, API Agents; OpenAI, Safety Measurement; Glean, AI Quality; Abridge, AI/ML (Evals); Scale AI, Forward Deployed PM, Enterprise; Decagon, Senior Agent PM; Sierra, Agent Development; Anthropic, Research Product Manager, Labs.
  3. Hamel Husain and Shreya Shankar, AI Evals: Everything You Need to Know (FAQ), updated September 2026.
  4. Lenny Rachitsky, OpenAI's CPO on how AI changes must-have skills (Kevin Weil), Lenny's Podcast, 10 April 2025.
  5. OpenAI, A practical guide to building agents.
  6. Microsoft, Guidelines for Human-AI Interaction (HAX Toolkit).
PM
Written by the AllthingsPM team
Frameworks and interview prep for product managers.
The AI PM course

Reading is the easy half.
The course grades the other half.

Start for free