Flash sale 30% off with code LAUNCH30 Ends in --:--:--
See pricing
All Things PM

What a Day Looks Like for an AI PM

An AI product manager spends the day on classic PM work plus four AI jobs: reading model outputs, running evals, deciding how much an agent may do on its own, and getting the product through a customer's security review. AllthingsPM teaches each one in an AI PM course built from 604 real PM job postings.

AllthingsPM·September 29, 2026·16 min read
A product manager at a desk reading a long printed transcript with a pencil, marking some lines, while a laptop and a small stack of test cards sit beside them
Half of the day is the classic PM job. The other half is reading what the model actually did.

An AI product manager does the whole PM job (find the problem, decide what to build, ship it, prove it worked) plus four things that fill a big share of the day: reading real model outputs, running evals that decide whether a change ships, choosing how much an agent is allowed to do on its own, and getting the product through a customer's security and deployment review. AllthingsPM is an AI PM course and PM interview prep platform, and its course is built from 604 real PM job postings, so the day below is assembled from what employers actually write in those postings, not from guesswork.

The rest of this post walks through that day hour by hour, shows how often each kind of work appears in real postings, and points you to the exact lesson that teaches it.

What does an AI product manager actually do all day?

Here is a composite day for a PM on an AI-native product, for example an assistant or agent sold to businesses. The hours are illustrative; the tasks are all taken from real job postings and the sources listed at the end.

TimeWhat the AI PM is doingWhere it comes fromLearn it on AllthingsPM
9:00Reads a sample of yesterday's conversations and flags failure patternsHamel Husain: "You can never stop looking at data"Failure attribution
10:00Eval review with engineering and research: did the new prompt or model clear the bar?Braintrust: PMs "decide on success criteria" and "analyze results"Evals chapter
11:00Customer call about a pilot, plus follow-ups from the security review79% of AI-native postings ask for enterprise deploymentShip it into somebody else's company
12:30Writes the spec for a new agent step, including where a human must approve73% of AI-native postings ask for agentsWorkflow or agent
14:00Builds a rough prototype to test an idea before asking for engineering timeAnthropic's Claude Science PM posting: "Prototype ideas yourself with Claude"Prototyping tools
15:30Pulls the numbers: task success, cost per call, retention62% of AI-native postings ask for outcomes and metricsSQL for PMs
16:30Roadmap, stakeholder updates, launch readiness criteria88% ask for classic PM craftThe AI PM role today

Two things stand out. First, most of the calendar still looks like any PM's calendar: customers, specs, roadmap, updates. The Institute of AI PM puts the classic share at about 60% of the job, with model evaluation, prompt work, data quality and AI UX as the rest. Second, the AI part is not a separate block. It leaks into every meeting, because the product can give a different answer to the same input and is wrong some of the time.

AllthingsPM (us) band first, then a bar chart of AI-native PM postings: technical fluency 90%, PM craft 88%, enterprise deployment 79%, AI UX and oversight 78%, agents 73%, 0 to 1 work 68%, outcomes and metrics 62%, evals 53%

The chart is the same day seen from the hiring side. We read 604 PM postings from 95 companies straight from their careers boards; 286 were AI-native roles. Each bar is the share of those 286 that ask for that kind of work. The full data is in our AI PM hiring report.

Why does an AI PM start the day reading transcripts?

Because a dashboard can tell you that task success dropped, but not why. In a normal product, a bug is reproducible: same input, same wrong output, fix it once. In an AI product, a failure is a pattern spread across many conversations, and fixing one case can break another.

So the first hour often goes to reading. Hamel Husain, who has helped many teams build AI products, writes that unsuccessful ones "almost always share a common root cause: a failure to create robust evaluation systems," and that "you can never stop looking at data." The Institute of AI PM makes the same point about dashboards: traditional PMs check them weekly, AI PMs check them daily, because AI products "can degrade in ways that traditional software can't."

What the PM is actually doing while reading:

  • Grouping failures into a small set of named types (wrong fact, ignored instruction, wrong tool, too slow).
  • Deciding which layer caused each one: the model, the context it was given, the code around it, or the screen the user saw.
  • Picking the one or two patterns worth fixing this week.

That last step is product judgment, not engineering. It is the same prioritization a PM has always done, applied to a new kind of evidence.

How AllthingsPM does this. The free foundations chapter of the AI PM course has a lesson on attributing every failure to a layer: model, context, harness or surface. You practice sorting real failure examples before you ever need to do it on the job.

What happens in an AI PM's eval review?

An eval is a repeatable test of whether the AI does the job well enough. The eval review is where a change gets a yes or a no. Braintrust, which makes eval software, lists the PM's part plainly: "PMs decide on success criteria, develop hypotheses about what to improve, label sample data, get real-world inputs into the eval platform, and analyze results."

In practice the meeting looks like this. Engineering has tried a new prompt, a new model or a new retrieval setup. The team looks at the pass rate on a fixed set of test cases, then at the cases that got worse. The PM asks the questions only the PM can answer: which failures matter most to users, whether a small quality gain is worth extra latency or cost, and whether the bar set before the work started has been cleared.

This is where "definition of done" changes. In classic software, done means acceptance criteria pass in QA. In AI products, done means an eval score clears a threshold everyone agreed on before the build. Anthropic's posting for a PM on Claude Science asks the PM to "define target behaviors, build and shape evals," and its Model Behaviors research PM posting asks for someone to "contribute to evals that measure alignment progress." Evals appear in 53% of AI-native postings in our corpus, and the literal word "eval" is about 13 times more common in AI-native postings than in others (32% against 3%).

How AllthingsPM does this. The evals chapter teaches how to define "good" before the build and make the number defensible, ending in a graded case study. For a deeper read, see AI evals for product managers, and look at a live evals product lead role at Abridge to see how one company writes the job.

Why do AI PMs spend so much time with customers' IT and security teams?

Because most AI products today are sold to businesses, and a business will not let an assistant or agent touch its data until security, legal and IT sign off. Enterprise and deployment work shows up in 79% of AI-native postings in our corpus. The Claude Science posting lists it directly: "Drive enterprise readiness: security and compliance reviews, admin and deployment controls."

A mid-morning customer call might cover:

  • Which data the product can see, and who in the customer's company can switch it on.
  • Admin controls: single sign-on, permissions, audit logs.
  • What happens when the AI is wrong on that customer's data, and who is told.
  • A success metric for the pilot that both sides agree on.

Some companies now hire "forward deployed" PMs who sit with the customer and get the agent into production. In our corpus, 21 AI-native postings used that phrase, and no other postings did.

How AllthingsPM does this. A full chapter, ship it into somebody else's company, covers security review, permissions and pilots. When you prepare for a role like this, paste the posting into the JD mock interview and practice the deployment questions that role will actually ask.

How does an AI PM decide what an agent is allowed to do?

This is the most distinctly AI part of the day, and the one that separates AI PM postings from the rest: 73% of AI-native postings ask for agent work against 34% of other PM postings.

The core decision is how much freedom to give the system. Anthropic's engineering guide draws the line clearly. Workflows are "systems where LLMs and tools are orchestrated through predefined code paths." Agents are "systems where LLMs dynamically direct their own processes and tool usage." Its advice: "we recommend finding the simplest solution possible, and only increasing complexity when needed."

So when an AI PM writes a spec for a new agent step, the spec answers questions a classic PRD never had to:

  • Could this be a fixed workflow instead of an open-ended agent?
  • Which actions need a human to approve them first, and what does the approval screen show?
  • What does the user see when the agent is unsure, slow or wrong?
  • What is the cost per successful task, not just per call?

Design for a product that is sometimes wrong is its own skill. AI UX and human oversight show up in 78% of AI-native postings.

How AllthingsPM does this. Start with the lesson workflow or agent, then the approval gate, which covers where a human check belongs and what it must show. Both chapters end in graded case studies built from the same postings.

Do AI PMs write code or build prototypes?

Many now build rough prototypes themselves, though they are not expected to ship production code. Anthropic's Claude Science posting asks the PM to "Prototype ideas yourself with Claude to validate them before committing engineering time." In our corpus, "prototype" appears in 15% of AI-native postings against 9% of others, and Claude Code or Cursor is named in 9% against 1%.

The afternoon prototype is usually small: a script that calls the model API with a draft prompt, a clickable demo, or a quick test of whether a model can handle a task at all. The point is to learn cheaply whether an idea works before a team spends a sprint on it.

How AllthingsPM does this. The PM as builder chapter includes prototype by what it must prove and build and iterate in Claude Code, Cursor and Codex. The free foundations chapter starts with making the API call yourself, which is the fastest way to lose any fear of the technical side.

Which numbers does an AI PM track?

The same outcome numbers any PM tracks (activation, retention, revenue), plus a few that only AI products have. Outcomes and metrics appear in 62% of AI-native postings, and model economics (cost and latency) in 20%.

The AI-specific numbers are usually:

  • Task success rate: did the user get what they came for.
  • Eval pass rate: the offline score from the eval review.
  • Cost per successful task: every model call costs money, unlike a normal feature that costs close to nothing once built.
  • Latency: how long the user waits.
  • Escalation or override rate: how often a human had to step in.

A new model version from a vendor can move all of these overnight, which is why the daily check matters.

How AllthingsPM does this. SQL for PMs teaches you to pull the number yourself, and the chapter prove it paid off covers outcomes, unit economics and pricing. Our guide to AI product metrics goes deeper on each number.

How is an AI PM's day different from a regular PM's day?

The first half of the answer is: less than people think. PM craft (PRDs, roadmap, strategy) appears in 88% of AI-native postings and 94% of other PM postings. Stakeholders, customers and prioritization still fill most of the calendar.

The difference is in what counts as evidence and what counts as done:

Part of the dayRegular PMAI PM
Morning checkWeekly dashboard reviewDaily read of real outputs and quality numbers
Definition of doneAcceptance criteria pass in QAEval pass rate clears an agreed bar
BugsReproduce, fix oncePatterns across many transcripts
SpecsFeatures and flowsAlso autonomy, approvals and failure states
CostNear zero per user once builtA cost on every model call
PartnersEngineering, design, dataPlus research, safety and the customer's IT team

If you already work as a PM, the AI layer is learnable on top of the skills you have. Our post on AI PM vs PM covers the differences in more detail, and how to become an AI PM with no experience covers the route in if you do not.

How AllthingsPM does this. The free lesson the AI PM job now explains the role through selection, taste and verification, and shows how a senior PM adds AI depth. It is the best single place to start if this post described a job you want.

Why AllthingsPM is the better choice for learning the AI PM job

The fastest way to understand an AI PM's day is to practice each part of it against what employers actually ask for. That is how AllthingsPM is built. The AI PM course comes from 604 real PM job postings, so the chapters line up with the day above: failure attribution, evals, agents, AI UX and human oversight, enterprise deployment, prototyping and outcomes. Each chapter ends in a graded case study, and the course is updated weekly as postings change.

Around the course sits everything you need to get the job. The jobs catalog holds 116 live PM job descriptions at 18 AI companies, each with its own mock. The question bank has 4,122 real questions from 260 companies, each with an answer guide. The JD mock interview builds a scored mock from any posting you paste, and resume review against a JD checks your resume against the same posting. There is also Resume Job Match for finding roles and 455 PM portfolios for examples of proof of work.

Other options have real strengths: cohort courses on Maven and Reforge give you live instructors and a peer group. For learning the day-to-day AI PM job and practicing for real roles in one place, with a free tier and Pro at $20 a month or $120 a year, AllthingsPM gives you more per dollar. Start with the free foundations chapter of the AI PM course.

Frequently asked questions

What does an AI product manager do?

An AI product manager does the full PM job and adds four AI tasks: reading real model outputs, running evals that decide whether a change ships, deciding how much an agent may do on its own, and getting the product through customers' security and deployment reviews. In our corpus of 286 AI-native PM postings, 73% ask for agent work and 53% for evals.

What is the best way to learn what an AI PM does day to day?

AllthingsPM is the best place to start: its AI PM course is built from 604 real PM job postings, so each lesson matches a real part of the day, and you can practice with mocks built from live AI company job descriptions. Pair it with reading a few real postings in full.

Do AI product managers need to code?

Not to ship production code, but many prototype. Anthropic's Claude Science PM posting asks the PM to prototype ideas with Claude before committing engineering time, and 90% of AI-native postings in our corpus ask for technical fluency. Being able to call a model API and read what comes back is enough to start.

How much of an AI PM's job is different from a regular PM's?

Most of it is the same. The Institute of AI PM estimates about 60% is classic PM work, and PM craft appears in 88% of AI-native postings in our corpus. The difference is in evidence (reading outputs and evals) and in new design decisions like agent autonomy and approvals.

What meetings does an AI PM have?

Typically an eval review with engineering and research, customer calls about pilots and security reviews, spec reviews for new agent steps, and the usual roadmap and stakeholder updates. Many also block time to read transcripts and to prototype.

Is AI product management a good career?

Demand is strong: 286 of the 604 PM postings we read (47%) were AI-native roles, at 78 of 95 companies. Entry-level roles are rare, though; only 1% of postings were APM level, so most people move in from an existing PM or adjacent role.

Ready to see if this is the job for you? Start the AllthingsPM AI PM course free and work through the foundations chapter this week.

Sources

  1. AllthingsPM JD corpus: 604 PM postings from 95 companies, read 6 September 2026. AI PM hiring 2026 report
  2. Anthropic, Product Manager, Claude Science (job posting). job-boards.greenhouse.io
  3. Anthropic, Research Product Manager, Model Behaviors (job posting). jobs.generalcatalyst.com
  4. Hamel Husain, "Your AI Product Needs Evals." hamel.dev
  5. Braintrust, "Evals for PMs: A practical guide to AI product quality." braintrust.dev
  6. Anthropic, "Building effective agents." anthropic.com
  7. Institute of AI PM, "A Day in the Life of an AI Product Manager." institutepm.com
  8. AllthingsPM pricing and plans (free tier, Pro $20 a month or $120 a year). AllthingsPM
PM
Written by the AllthingsPM team
Frameworks and interview prep for product managers.
The AI PM course

Reading is the easy half.
The course grades the other half.

Start for free