Flash sale 30% off with code LAUNCH30 Ends in --:--:--
See pricing
All Things PM

AI Product Discovery: Interviews, Logs, and the Behavior You Can Move

AI product discovery combines story-based interviews with reading conversation logs, tickets and traces, then roots every idea in one user behavior you can move. AllthingsPM teaches it as a full course chapter with a graded case.

AllthingsPM·September 29, 2026·17 min read
A product manager at a whiteboard lays out a long row of small paper tiles, each tile one word fragment, while a narrow window frame on the wall shows only part of the row
Interviews tell you the story. Logs tell you what people actually typed.

AI product discovery is the work of finding which user problem an AI feature should solve before you build it, using two sources together: story-based interviews (what people did the last time they faced the problem) and logs (conversation transcripts, support tickets and traces that show what they actually asked for). The goal is not a list of ideas. It is one user behavior you can measurably move.

AllthingsPM is an AI PM course and PM interview prep platform. Its course has a full chapter on this skill, Discovery and strategy for AI products, and it opens with the lesson this post is named after: Discovery for AI: interviews, logs, and the one behavior you can actually move. The course is built from 604 real PM job postings and updated weekly.

What is AI product discovery, and how is it different from normal discovery?

Classic discovery, as Teresa Torres describes it, means interviewing customers every week, mapping what you hear onto an opportunity solution tree that starts from a clear outcome, and testing assumptions before you build [1][2]. None of that goes away with AI. Three things change.

  1. The capability is uncertain. With normal software you know what you can build. With a model, you often do not know if it can do the task well enough until you try it on real cases.
  2. You have a new, richer data source. Once an AI feature is live, users type their intent in their own words. Every conversation log is a free, unprompted statement of what someone wanted.
  3. Averages lie harder. A dashboard can show a healthy thumbs-up rate while one important intent fails every time. Aggregate metrics hide the failure; reading the actual conversations finds it.

So AI product discovery is classic discovery plus log reading plus an early feasibility test, all pointed at one behavior.

Discovery sourceWhat it tells youWhat it missesWhen to use it
Story-based interviewsContext, motivation, the workaround, why they switchedScale, exact phrasing, how oftenBefore any build, and weekly after
Conversation logsReal intent in the user's words, unexpected usesWhy they asked, what they did nextAs soon as any AI surface is live
Support ticketsFailures users cared enough to reportSilent failures, people who just leftAlways; the cheapest labeled data you have
TracesWhich step failed: retrieval, tool call, model, UIUser motivationWhen a log shows a failure you must explain
Capability probesWhether the model can do the job on real casesWhether anyone wants itBefore you commit a roadmap slot

How AllthingsPM does this: the customer interviews lesson teaches the interview and the log channels as one practice, and the next lessons cover the assumption test that could kill the idea and how to qualify the opportunity before it reaches a roadmap.

Why do AI discovery interviews go wrong?

The classic failure is the opinion interview. You ask "Would you use an AI assistant that summarized your meetings?" and people say yes, because saying yes is polite and costs nothing. You leave with a feature request, not evidence.

AI makes this worse. The word "AI" invites people to describe an imagined magic product. Nobody knows what they would do with a capability they have never had, so their predictions are guesses.

The fix is the story-based interview: ask about a specific, recent instance. Torres recommends collecting stories about past behavior rather than asking people to speculate [1][2].

Compare the two:

  • Opinion question: "How would AI help you write status reports?"
  • Story question: "Tell me about the last status report you wrote. When did you start? What did you open first? Where did you get stuck?"

The second one tells you where the time actually goes, which tools are open, what gets copied from where, and which step the person hates. That is where an AI feature can earn a place.

How AllthingsPM does this: the discovery lesson drills the "last time you did it" question and the opinion interview failure mode. If you want the source material, read the Continuous Discovery Habits summary or our comparison of Continuous Discovery Habits vs The Mom Test.

Who should you interview for an AI product?

Recruit people who recently switched. Someone who changed how they do a job in the last few weeks can still remember the moment clearly. That is the core of the Jobs to Be Done switch interview developed by Bob Moesta and Chris Spiek [3].

Their four forces model is useful for AI products because adoption stalls on anxiety far more than on the feature list [3]:

  • Push: what made the old way intolerable ("I spent every Friday afternoon on this report").
  • Pull: what attracted them to the new way ("a teammate showed me a draft it wrote in a minute").
  • Anxiety: what could go wrong ("what if it makes up a number and my boss sees it").
  • Habit: the comfort of the old way ("I know my spreadsheet").

A switch happens only when push plus pull beats anxiety plus habit [3]. For AI features, anxiety is usually about being wrong in front of someone. That tells you a lot about the product: show sources, make edits easy, and keep a manual path.

Also interview the adjacent user: the person who receives the output. The manager who reads the AI-drafted report is part of the job, and their trust decides whether the drafter keeps using it.

Segment by circumstance, not demographics. "Sales reps preparing for a first call with a new account in under 15 minutes" is a segment you can design for. "Mid-market sales reps aged 25 to 34" is not.

How do you use logs, tickets and traces for discovery?

This is the part most teams skip, and it is the most valuable once anything is live. The AllthingsPM course calls these "the discovery channels your competitors ignore", and its demand research notes that the discovery move that works when there are no users yet is reading logs, tickets and traces as the cheapest labeled data available.

Conversation logs

A conversation log is a user stating their intent, unprompted, in their own words. Reading a sample of them answers questions no survey can:

  • What do people actually ask for, compared with what we designed for?
  • Which intents keep coming back?
  • What are people using this for that we never planned? Unintended usage is a signal, sometimes the biggest one.

Anthropic's Clio is a public example of this at scale. It analysed 1 million claude.ai conversations and found that web and mobile app development made up over 10 percent of conversations, education more than 7 percent and business strategy nearly 6 percent, along with thousands of niche uses [4]. It did this with privacy built in: Claude extracted facets, clustered similar conversations and wrote cluster summaries that omit private details, with minimum thresholds so rare topics are not exposed, and human analysts saw only aggregated summaries, not raw conversations [4][5].

You do not need Clio. You need a sample, a spreadsheet and a privacy rule agreed with legal before anyone reads a transcript.

Support tickets

Tickets are failures that users cared enough to report. They are already labeled with the user's frustration. The course makes one rule here non-negotiable: a ticket without a trace id is unactionable. If support cannot link a complaint to the exact run that produced it, engineering cannot tell whether the model, the retrieval or the interface failed.

Traces

When a log shows a failure, the trace shows where it happened. The method for reading them comes from AI evals practice: sample real traces, write an open note on the first thing that went wrong in each (open coding), then group the notes into a failure taxonomy (axial coding) and count them [6][7]. Hamel Husain suggests annotating at least 30 traces yourself before letting an LLM propose groupings, and reviewing around 100 diverse traces overall [6][7].

That count is discovery output. A ranked table of failure types is a ranked table of opportunities.

How AllthingsPM does this: the course teaches the log and ticket channels in the discovery lesson, then turns the same traces into a ranked failure backlog and an eval suite in the evals chapter. For the eval side in full, read AI evals for product managers.

Bar chart: AllthingsPM (us) teaches discovery as a full course chapter; AI PM craft appears in 89 percent of 335 AI PM postings, 0-to-1 under ambiguity in 69 percent, evals in 53 percent
AllthingsPM course demand report: 335 live AI-native PM postings across 88 companies, September 2026

What does "the behavior you can move" mean?

This is the idea that holds the whole practice together. Your opportunity solution tree needs a root, and Torres puts a clear outcome at the top [1]. The course sharpens this for AI: root the tree in a user behavior you can move, not in a business number you cannot directly touch.

Here is the difference:

  • Business outcome: "Increase net revenue retention." True, important, and no single feature moves it in a quarter.
  • Product outcome: "Increase the share of account managers who send a renewal summary within 24 hours of a customer call." You can observe it, measure it weekly and build for it.

Once the root is a behavior, every branch has to earn its place by explaining why that behavior is not happening today. Your interview stories and log clusters become the opportunities. The AI feature is one solution among at least three you consider for each opportunity, and sometimes the winning solution is not AI at all.

Three rules keep this honest:

  1. Size the opportunity without thinking about effort. Decide how much it matters first, then how hard it is.
  2. Generate three solutions per opportunity. If the only one on the page is "add an AI assistant", you have not done discovery.
  3. Pick the behavior before launch, so you can measure it after. This connects straight to AI product metrics, where task success is the number that proves the feature paid off.

How AllthingsPM does this: the discovery lesson teaches product outcome versus business outcome and opportunity sizing, then the strategy memo lesson and the roadmap planning lesson turn the chosen behavior into a plan around an evaluable slice.

How do you test feasibility during AI discovery?

With AI, "can we build it" is a real question, so it belongs in discovery, not after it. The cheapest test is a capability probe: take 20 to 50 real examples from your logs or interviews, run them through the model with a simple prompt, and read every output.

You are looking for one of three answers:

  • It works on most cases. Move on to desirability and cost.
  • It works on some cases. Find the pattern. Often a narrower segment is the real first product.
  • It fails in ways a better prompt will not fix. Stop, or change the problem.

A second cheap test is the Wizard of Oz: a human does the AI's job behind the interface for a few users. It tells you whether people want the output before anyone builds the model pipeline.

Then write a falsifiable problem statement: one sentence a single test could prove wrong. "Account managers skip renewal summaries because drafting one takes over 20 minutes" can be proven wrong by timing five people. "Users want AI help with renewals" cannot.

How AllthingsPM does this: the assumption testing lesson covers assumption mapping, the riskiest assumption test, prompt experiments as feasibility evidence and the falsifiable statement, and the qualify the opportunity lesson asks whether an agent is even the right answer.

Can you use AI to do the discovery synthesis?

Yes, with a rule: every insight must trace back to a source. AI is good at clustering hundreds of notes or transcripts quickly. It is also able to produce a confident theme nobody said.

Nielsen Norman Group's guidance is a good line to hold. Synthetic users, AI-generated profiles that simulate a segment, can help with desk research and generating hypotheses, but user research needs real users and synthetic ones should complement, not replace, real research [8].

A practical workflow:

  1. Code the first 30 notes or transcripts yourself so you know what the categories should look like [6].
  2. Let a model propose clusters for the rest.
  3. For each cluster, open three raw quotes or log lines behind it. If you cannot find them, delete the cluster.
  4. Keep the link from every insight to its source in the synthesis document.

How AllthingsPM does this: the discovery lesson covers AI-assisted synthesis, source traceability and fabrication risk as part of the same practice, so the verification step is taught with the shortcut, not after it. To see how these ideas connect to the rest of the AI PM toolkit, browse the AI PM knowledge graph.

How does AI product discovery show up in PM interviews?

AI labs ask about it directly. The AllthingsPM question bank includes real prompts such as "Design a discovery process for Claude Science" and one about three inputs on Claude Code performance: user interviews, production transcripts and more, which is exactly the interviews-versus-logs tradeoff in this post. Another asks how PMs at Anthropic should uncover non-obvious use cases for frontier models.

A strong answer has the same shape as good discovery:

  1. Name the behavior you want to move.
  2. Say which sources you would use and why: interviews for the why, logs for the what and how often, traces for where it breaks.
  3. Explain how you would protect user privacy while reading logs.
  4. Describe the one test that could kill your leading idea.
  5. Say what you would measure after launch.

How AllthingsPM does this: every one of the 4,122 questions in the question bank has its own page and answer guide, and you can practise any of them out loud in a scored mock interview, in text or voice, with follow-ups.

Why AllthingsPM is the better choice for learning AI product discovery

Most discovery material was written before AI features existed. Books like Continuous Discovery Habits and The Mom Test are still the right foundation, and AllthingsPM has summaries of the key discovery books so you can take the core ideas in an evening. What they do not cover is the AI layer: log reading, ticket and trace channels, capability probes and verified AI synthesis.

General AI PM courses tend to go the other way and spend their time on models and evals. The AllthingsPM demand report on 335 live AI-native PM postings found classic PM craft in 89 percent of postings and building under ambiguity in 69 percent, ahead of evals and measurement at 53 percent. Employers want PMs who can find the right problem, not only PMs who can test a model.

AllthingsPM teaches both halves in one place. The Discovery and strategy for AI products chapter takes you from the interview to the log to the assumption test to the strategy memo, and ends in a graded case on a product you carry through the whole course. Then you practise the same skill against real interview questions from AI labs, in mocks built from real job descriptions, and you can check your resume against a target role with the resume review.

Book summaries and coaching programs have their place; for learning AI discovery as a repeatable practice and proving it in an interview, AllthingsPM gives you the chapter, the questions and the mock at $20 a month or $120 a year, with a free tier to start. Open the discovery chapter.

Frequently asked questions

What is AI product discovery?

AI product discovery is finding which user problem an AI feature should solve before building it. It combines story-based interviews with reading conversation logs, support tickets and traces, and it roots every idea in one user behavior you can measurably move. It also includes an early test of whether the model can do the job.

What is the best way to learn AI product discovery?

AllthingsPM is the best place to start: its AI PM course has a full chapter on discovery and strategy for AI products, from interviews and logs to assumption tests and a graded strategy memo case, built from 604 real PM job postings. Pair it with the Continuous Discovery Habits summary for the classic foundation.

How many customer interviews should a PM do for an AI feature?

Teresa Torres recommends that product trios interview at least one customer every week, as an ongoing habit rather than a one-off project [2]. For AI features, add a regular log review, because logs show you intent at a scale interviews cannot.

Can I read user conversation logs for discovery?

Only under a privacy policy agreed with your legal team, with access controls and redaction. Anthropic's Clio shows one approach: automated clustering and summaries that omit private details, with analysts seeing only aggregated results [4][5].

Should I use synthetic users instead of interviews?

No. Nielsen Norman Group says synthetic users can help with desk research and hypotheses, but should complement, not replace, research with real users [8]. Use them to prepare questions, never to make the decision.

How is AI discovery different from AI evals?

Discovery decides which problem to solve and which behavior to move. Evals check whether the model's output is good enough. They share a method, reading real traces and naming failures, which is why the AllthingsPM course connects the discovery chapter to the evals chapter.

Start learning AI product discovery today

Open the AllthingsPM course, start with the interviews and logs lesson, and practise one discovery question in a free mock interview the same day.

Sources

  1. Product Talk, "Opportunity Solution Trees: Visualize Your Discovery to Stay Aligned and Drive Outcomes": https://www.producttalk.org/opportunity-solution-trees/
  2. Lenny's Newsletter, "Teresa Torres on how to interview customers": https://www.lennysnewsletter.com/p/teresa-torres-on-how-to-interview
  3. JobsToBeDone.org, "The Four Forces of Progress: Push, Pull, Anxiety and Habit": https://jobstobedone.org/the-four-forces/
  4. Anthropic, "Clio: Privacy-preserving insights into real-world AI use": https://www.anthropic.com/research/clio
  5. Tamkin et al., "Clio: Privacy-Preserving Insights into Real-World AI Use", arXiv 2412.13678: https://arxiv.org/abs/2412.13678
  6. Hamel Husain, "Why is error analysis so important in AI evals, and how is it performed?": https://hamel.dev/blog/posts/evals-faq/why-is-error-analysis-so-important-in-llm-evals-and-how-is-it-performed.html
  7. Hamel Husain, "AI Evals: Everything You Need to Know": https://hamel.dev/blog/posts/evals-faq/
  8. Nielsen Norman Group, "Synthetic Users: If, When, and How to Use AI-Generated Research": https://www.nngroup.com/articles/synthetic-users/
  9. AllthingsPM course demand report, 335 live AI-native PM postings across 88 companies, September 2026: https://allthingspm.app/course
PM
Written by the AllthingsPM team
Frameworks and interview prep for product managers.
The AI PM course

Reading is the easy half.
The course grades the other half.

Start for free