Flash sale 30% off with code LAUNCH30 Ends in --:--:--
See pricing
All Things PM

AI PM Glossary: 200 Terms Every AI PM Should Know

A 200-term AI product management glossary in plain English, grouped into ten areas from tokens and RAG to evals, agents, safety and pricing. AllthingsPM's AI PM course teaches every term in context, with 212 linked concepts in its knowledge graph.

AllthingsPM·September 29, 2026·22 min read
A product manager at a long table sorting a tall stack of printed job descriptions into labelled piles with a highlighter, sticky notes marking repeated phrases
Two hundred words that come up in every AI PM interview, spec review and model launch.

An AI product management glossary is only useful if each term tells you what decision it changes. This one does: 200 terms in ten groups, each defined in one line from a PM's point of view. The fastest way to go from knowing the words to using them is AllthingsPM, whose AI PM course is built from 604 real PM job postings and whose knowledge graph maps 212 AI concepts and 1,113 connections between them.

AllthingsPM is an AI PM course and PM interview prep platform. Every group below ends with the exact lesson or tool where the terms are taught and practiced.

Bar chart of AI PM terms covered: AllthingsPM knowledge graph first with 212 concepts, this AllthingsPM glossary with 200 terms, then The Product Space 120, AI Product Craft 100, Product Leadership 80 plus and a Medium glossary 20 plus

What are the core model terms every AI PM should know?

TermWhat it means for a PM
Artificial intelligence (AI)Software that performs tasks we associate with human judgment.
Machine learning (ML)Systems that learn patterns from data instead of hand-written rules.
Deep learningML using many-layered neural networks.
Neural networkLayers of weighted connections that turn inputs into outputs.
Generative AIModels that produce new text, images, audio or code.
Large language model (LLM)A model with many parameters trained on vast text to generate language [1].
Foundation modelA large general model that many products build on.
Frontier modelThe most capable model a lab currently ships.
TransformerThe attention-based architecture behind modern LLMs.
AttentionHow a model weighs which earlier tokens matter for the next one.
ParametersThe learned weights; a rough proxy for model size.
TokenThe unit a model reads and writes; for Claude about 3.5 English characters [1].
TokenizerThe code that splits text into tokens.
Context windowThe model's working memory for one request [1].
InferenceRunning a trained model to get an output [3].
Next-token predictionThe core task: guess the most likely next token.
TemperatureA setting for randomness; lower is more predictable [1].
Non-determinismThe same input can give different outputs, even at temperature 0 [1].
Reasoning modelA model that spends extra tokens thinking before it answers.
Multimodal modelA model that handles more than one input type, such as text plus images.

How AllthingsPM does this. The free Foundations chapter opens with context windows and attention and pretraining, post-training and reasoning models, and explains why you cannot promise determinism. Each term links into the knowledge graph, so you see what it depends on.

Which prompting and context terms matter most?

TermWhat it means for a PM
PromptThe input you send the model.
System promptStanding instructions that shape every response.
User promptThe part of the input written by the end user.
Prompt engineeringWriting and testing prompts to get reliable behavior.
Context engineeringDeciding what information goes into the window, and in what order.
Zero-shot promptingAsking for a task with no examples.
Few-shot promptingIncluding a few worked examples in the prompt [3].
Chain-of-thoughtAsking the model to reason step by step [3].
Prompt templateA reusable prompt with slots for variables.
Prompt chainingSplitting a task into several prompts run in sequence.
Structured outputForcing replies into a schema such as JSON.
JSON modeA setting that guarantees parseable JSON.
Stop sequenceText that tells the model to stop generating.
Max tokensThe cap on output length, and on output cost.
PrefillStarting the model's reply for it to steer format.
PersonaA role the prompt asks the model to adopt.
GroundingTying answers to supplied sources rather than memory.
Prompt cachingReusing a repeated prompt prefix to cut cost and latency.
Context rotQuality dropping as the window fills with noise.
CompactionSummarizing old context so a long task fits the window.

How AllthingsPM does this. The lesson Context as a budget treats the system prompt as a product surface and shows when retrieval, not a longer prompt, is the fix. You then practice explaining that trade-off in a mock interview.

What do RAG and retrieval terms mean?

TermWhat it means for a PM
Retrieval-augmented generation (RAG)Fetching relevant documents at runtime and passing them to the model [1].
Knowledge baseThe document set RAG searches.
EmbeddingA vector that captures meaning, so similar items sit close together [3].
VectorA list of numbers representing text or images.
Vector databaseA store built for fast similarity search over vectors.
Semantic searchSearch by meaning rather than exact words.
Keyword search (BM25)Classic word-match ranking.
Hybrid searchCombining semantic and keyword search.
ChunkingSplitting documents into retrievable pieces.
Chunk sizeHow big each piece is; too big dilutes, too small loses context.
Top-kHow many results retrieval returns.
RerankerA second model that reorders retrieved results.
Retrieval recallShare of relevant documents actually retrieved.
Retrieval precisionShare of retrieved documents that are relevant.
CitationShowing which source supports each claim.
Metadata filteringNarrowing retrieval by date, owner or permissions.
Index freshnessHow current the searchable data is.
Agentic searchA model deciding what to search for, iteratively.
Deep researchMulti-step agentic search that produces a report.
Knowledge graphEntities and their relationships stored as a network.

How AllthingsPM does this. Retrieval intuition and the agentic search and Deep Research patterns are taught in chapter 6, with cost trade-offs spelled out. The AllthingsPM knowledge graph is itself a working example of the last term in this table.

Which AI agent terms come up in interviews?

TermWhat it means for a PM
AI agentA model that directs its own steps and tool use toward a goal [4].
WorkflowModel calls orchestrated through predefined code paths [4].
AgenticDescribes systems with some autonomy over their next step.
Tool use (function calling)The model requesting an external action, such as an API call.
Tool schemaThe description that tells the model what a tool does and needs.
HarnessThe loop, retries and stop conditions wrapped around the model.
Agent loopThink, act, observe, repeat until done.
Stop conditionThe rule that ends the loop.
PlanningThe agent breaking a goal into steps.
MemoryWhat the agent keeps between steps or sessions.
OrchestratorThe component that routes work between models or agents.
SubagentA helper agent given one slice of the task.
Multi-agent systemSeveral agents collaborating on one goal.
Model Context Protocol (MCP)An open standard for connecting AI apps to data, tools and workflows [2].
MCP serverA program that exposes tools or data over MCP [2].
MCP clientThe app side that connects to MCP servers [2].
Computer useAn agent operating a screen, mouse and keyboard.
Human in the loopA person approves or corrects agent actions.
Autonomy levelHow much the agent may do without asking.
Task completion rateShare of agent runs that finish the job correctly.

How AllthingsPM does this. Chapter 6, Agents and agentic architecture, covers what is inside the harness and when multi-agent is worth it. For a shorter read, see agents vs workflows and the anatomy of an AI agent.

What training and customization terms should a PM understand?

TermWhat it means for a PM
PretrainingInitial training on a large unlabeled corpus to predict the next word [1].
Post-trainingEverything after pretraining that makes a model useful and safe.
Fine-tuningFurther training a pretrained model on additional data [1].
Supervised fine-tuning (SFT)Fine-tuning on example inputs and ideal outputs.
RLHFTraining from human rankings of outputs [1].
Reward modelA model that scores outputs to guide training.
Constitutional AITraining against written principles rather than only human labels.
DistillationA smaller model learning from a larger one [3].
QuantizationLowering weight precision to save memory and speed up inference [3].
LoRAA cheap fine-tuning method that trains small add-on weights.
Open-weight modelA model whose weights you can download and host.
Closed modelA model available only through the vendor's API.
Training dataThe examples the model learns from.
Synthetic dataTraining or test data generated by a model.
Data labelingHumans marking the correct answer on examples.
OverfittingLearning training data so well that new data suffers [3].
Mixture of expertsSpecialized sub-networks with a gate choosing which runs [3].
Model cardA document describing a model's uses, limits and evaluations.
Data flywheelProduction usage producing data that improves the next model.
Knowledge cutoffThe date after which the model has no training data.

How AllthingsPM does this. The course teaches the fix hierarchy first: prompt, context and harness before touching weights, a point made in the harness lesson. The data flywheel lesson shows how production turns into the next model.

Which eval and quality terms do AI PMs use daily?

TermWhat it means for a PM
EvalA repeatable test of whether the AI output is good enough.
Golden setA curated set of inputs with known good outputs.
Ground truthThe correct answer used for scoring [3].
Offline evalTesting before release on a fixed dataset.
Online evalMeasuring quality on live traffic.
LLM as a judgeUsing a model to grade another model's output.
RubricWritten criteria the grader scores against.
Human evalPeople grading outputs, the slowest and most trusted signal.
Pairwise comparisonPicking the better of two outputs.
BenchmarkA public standard test used to compare models.
Benchmark saturationTop models scoring so high a test stops separating them.
PrecisionCorrect positive predictions over all positive predictions [3].
RecallShare of actual positives the model catches [3].
F1 scoreThe balance of precision and recall.
HallucinationA confident answer that is false.
FaithfulnessWhether the answer matches the supplied sources.
Regression testChecking a change did not break what worked.
Error analysisReading failures and sorting them into causes.
Failure attributionDeciding which layer failed: model, context, harness or surface.
Evaluation awarenessA model behaving differently when it senses a test.

How AllthingsPM does this. Chapter 8, Evals: define good and make the number defensible, is one of the deepest in the course, and the failure attribution lesson teaches the four-layer diagnosis. Pair it with the AI evals guide for PMs and evaluation awareness.

What cost, latency and infrastructure terms matter?

TermWhat it means for a PM
LatencyTime between sending a prompt and getting the reply [1].
Time to first token (TTFT)How fast the first token appears [1].
Tokens per secondOutput speed once generation starts.
StreamingShowing output as it is generated.
Input tokensTokens you send; billed per token.
Output tokensTokens the model writes; usually priced higher.
Cost per taskTotal model spend to complete one user job.
Batch APIAsynchronous processing at a discount for non-urgent work.
Rate limitCap on requests or tokens per minute.
ThroughputHow much work the system handles at once.
GPUThe chip most models train and run on.
Model routingSending each request to the cheapest model that can handle it.
Fallback modelA backup model used when the first fails.
CachingStoring repeated results to save cost and time.
Edge inferenceRunning a model on the device.
Self-hostingRunning an open-weight model on your own servers.
APIThe interface your product uses to call the model.
SDKA code library that wraps the API.
ObservabilityLogging and tracing each model call.
TraceThe full record of one request, including tool calls.

How AllthingsPM does this. Chapter 2, Data fluency, teaches you to read logs and traces yourself, and chapter 9 covers unit economics and pricing. These are the numbers interviewers push on in AI product sense rounds.

What safety, security and governance terms should an AI PM know?

TermWhat it means for a PM
GuardrailA check that blocks or fixes unsafe input or output.
AlignmentMaking model behavior match intended goals and values.
HHHHelpful, honest, harmless: Anthropic's research framework [1].
Red teamingDeliberately trying to make the system fail.
JailbreakA prompt that bypasses safety behavior.
Prompt injectionInput that alters the model's instructions; OWASP's top LLM risk [5].
Indirect prompt injectionMalicious instructions hidden in retrieved content.
Sensitive information disclosureThe model leaking private data [5].
System prompt leakageUsers extracting hidden instructions [5].
Excessive agencyGiving an agent more power than the task needs [5].
Data poisoningCorrupting training, fine-tuning or embedding data [5].
Unbounded consumptionRunaway usage that drives cost or outages [5].
Content moderationFiltering harmful content.
BiasSystematic unfair differences in outputs across groups.
ExplainabilityBeing able to say why the model produced an output.
PIIPersonally identifiable information.
Data retentionHow long prompts and outputs are stored.
Audit logA tamper-evident record of who did what.
Responsible AIPractices for building AI that is fair, safe and accountable.
AI governancePolicies and owners for how AI is built and used.

How AllthingsPM does this. The AI PRD lesson makes you name risks, guardrails and success metrics before anything is built. Real questions such as how Claude Code should handle hundreds of parallel subagents test this area directly.

Which AI UX and product terms show up in specs?

TermWhat it means for a PM
CopilotAI that assists while the human stays in charge.
AutopilotAI that completes the task on its own.
Chat interfaceA conversational UI.
Ambient AIAI that acts in the background without being asked.
Suggested promptsStarter prompts that teach users what is possible.
Confidence indicatorUI that signals how sure the system is.
Graceful failureA useful response when the AI cannot help.
EscalationHanding a case to a human.
Feedback loopThumbs, edits and corrections captured as data.
UndoLetting users reverse an AI action.
Approval stepA human check before an agent acts.
Trust calibrationUsers trusting the AI as much as it deserves, no more.
OverrelianceUsers accepting wrong AI output without checking.
PersonalizationAdapting output to a specific user.
AI-native productA product that could not exist without the model.
WrapperA thin product on top of someone else's model.
MoatWhat stops rivals copying you, such as data or distribution.
Model dependencyRisk from relying on one vendor's model.
Build vs buyChoosing to train, fine-tune or call an API.
PrototypeA quick build to test whether the AI can do the job.

How AllthingsPM does this. Chapter 7, AI UX and human oversight, is about designing for a system that is wrong sometimes. The problem-first test lesson asks whether an agent is even the right answer.

What metrics, business and career terms complete the list?

TermWhat it means for a PM
AI product managerA PM who owns products whose core behavior comes from a model.
Agent product managerA PM who owns an AI agent's outcomes, often in customer service.
Forward deployed engineerAn engineer embedded with customers to ship AI into their systems.
AI PRDA spec that defines quality bars, risks and evals, not just features.
Success metricThe number that says the feature worked.
Guardrail metricA number that must not get worse.
Resolution rateShare of issues the AI solves without a human.
Deflection rateShare of contacts that never reach a human.
Acceptance rateShare of AI suggestions users keep.
A/B testComparing two versions on live users.
AdoptionShare of eligible users who use the feature.
RetentionShare of users who come back.
Gross marginRevenue left after model and serving costs.
Seat pricingCharging per user.
Usage pricingCharging per token, call or task.
Outcome pricingCharging per result, such as a resolved ticket.
Brownfield deploymentShipping AI into an existing company's messy systems.
PilotA limited rollout to prove value.
Product senseJudgment about what to build and why.
AI PM interview loopThe rounds AI PM candidates face, from product sense to technical.

How AllthingsPM does this. Chapter 9 teaches seat, usage and outcome pricing, chapter 10 covers brownfield enterprise deployment, and chapter 14, Get the job, walks the AI PM interview loop. Then practice on one of 116 live AI company job descriptions in the jobs catalog, such as the Product Lead, AI/ML (Evals) role at Abridge, or paste any JD into the JD mock interview.

How should you actually learn these 200 terms?

Learn them in the order a real product forces on you, which is the order of the ten tables above. Then test yourself out loud. An interviewer will not ask you to define RAG. They will ask why your support bot gives wrong refund answers, and expect chunking, retrieval recall, grounding and an eval in the same breath.

How AllthingsPM does this. Every course chapter ends in a graded case study. The question bank has 4,122 real questions from 260 companies, each with an answer guide, and any one can start a scored mock.

Why AllthingsPM is the better choice for learning AI PM vocabulary

Free glossaries from The Product Space (120 terms), AI Product Craft (100) and Product Leadership (80 plus) are clean references worth a bookmark [6][7][8]. But a list cannot tell you when your answer is wrong. AllthingsPM adds three things a glossary cannot:

  1. A course built from demand. 14 chapters and 101 lessons, derived from 604 real PM job postings and updated weekly, so the terms you learn are the ones employers ask for.
  2. A map of how ideas connect. The knowledge graph links 212 AI concepts through 1,113 connections, so you see that evals depend on ground truth, which depends on labeling.
  3. Practice that scores you. Mock interviews built from any job description, in text or voice, with follow-ups and a score, plus 4,122 real questions.

All of it sits in one account with a free tier (one JD mock a day) and Pro at $20 a month or $120 a year. For anyone serious about getting into AI product management, the verdict is simple: bookmark a glossary, but learn on AllthingsPM. Start the AI PM course free.

Frequently asked questions

What is the best AI product management glossary?

AllthingsPM's glossary is the best starting point because each of its 200 terms links to a lesson in an AI PM course built from 604 real job postings, and its knowledge graph maps 212 concepts. The Product Space and AI Product Craft offer useful free lists of 120 and 100 terms.

What AI terms should a product manager learn first?

Start with LLM, token, context window, RAG, embeddings, hallucination, evals, agents, guardrails and inference cost. These map directly to decisions about quality, cost and safety. The free Foundations chapter covers the first group.

What is the difference between an AI agent and a workflow?

Anthropic defines workflows as model calls orchestrated through predefined code paths, and agents as systems where the model dynamically directs its own process and tool use [4]. Workflows are more predictable; agents handle open-ended tasks.

Do AI PMs need to know how to code?

Most AI PM roles do not require production coding, but they do expect you to read logs, run basic SQL and reason about evals. AllthingsPM's Data fluency chapter teaches exactly that level.

What is an eval in AI product management?

An eval is a repeatable test of whether a model's output is good enough for your product, scored against a golden set, a rubric or human judgment. It is how AI PMs turn "it seems better" into a defensible number.

How do I use these terms in an interview?

Use them inside a diagnosis, not as definitions. Practice with a JD mock interview built from the role you want, and check your resume uses the right vocabulary with a resume review against the JD.

Ready to use the vocabulary?

Knowing 200 words is not the job. Making a trade-off with them is. Start the AllthingsPM AI PM course free and take your first scored mock the same day.

Sources

  1. Anthropic, Glossary, Claude documentation, accessed 29 September 2026.
  2. Model Context Protocol, What is the Model Context Protocol (MCP)?, accessed 29 September 2026.
  3. Google for Developers, Machine Learning Glossary, accessed 29 September 2026.
  4. Anthropic, Building effective agents, December 2024.
  5. OWASP GenAI Security Project, 2025 Top 10 for LLM Applications, accessed 29 September 2026.
  6. The Product Space, 120 AI Terms Every Product Manager Should Know in 2026.
  7. AI Product Craft, The Top 100 AI and ML Terms AI Product Managers Need to Know.
  8. Product Leadership, AI Product Management Glossary: 80+ Key Terms.
  9. AI Product Management Coaching, The Ultimate AI Product Management Glossary: 20+ Terms, Medium.
  10. AllthingsPM platform counts (chapters, lessons, concepts, connections, questions, job descriptions), September 2026.
PM
Written by the AllthingsPM team
Frameworks and interview prep for product managers.
The AI PM course

Reading is the easy half.
The course grades the other half.

Start for free