An AI product management glossary is only useful if each term tells you what decision it changes. This one does: 200 terms in ten groups, each defined in one line from a PM's point of view. The fastest way to go from knowing the words to using them is AllthingsPM, whose AI PM course is built from 604 real PM job postings and whose knowledge graph maps 212 AI concepts and 1,113 connections between them.
AllthingsPM is an AI PM course and PM interview prep platform. Every group below ends with the exact lesson or tool where the terms are taught and practiced.
What are the core model terms every AI PM should know?
| Term | What it means for a PM |
|---|---|
| Artificial intelligence (AI) | Software that performs tasks we associate with human judgment. |
| Machine learning (ML) | Systems that learn patterns from data instead of hand-written rules. |
| Deep learning | ML using many-layered neural networks. |
| Neural network | Layers of weighted connections that turn inputs into outputs. |
| Generative AI | Models that produce new text, images, audio or code. |
| Large language model (LLM) | A model with many parameters trained on vast text to generate language [1]. |
| Foundation model | A large general model that many products build on. |
| Frontier model | The most capable model a lab currently ships. |
| Transformer | The attention-based architecture behind modern LLMs. |
| Attention | How a model weighs which earlier tokens matter for the next one. |
| Parameters | The learned weights; a rough proxy for model size. |
| Token | The unit a model reads and writes; for Claude about 3.5 English characters [1]. |
| Tokenizer | The code that splits text into tokens. |
| Context window | The model's working memory for one request [1]. |
| Inference | Running a trained model to get an output [3]. |
| Next-token prediction | The core task: guess the most likely next token. |
| Temperature | A setting for randomness; lower is more predictable [1]. |
| Non-determinism | The same input can give different outputs, even at temperature 0 [1]. |
| Reasoning model | A model that spends extra tokens thinking before it answers. |
| Multimodal model | A model that handles more than one input type, such as text plus images. |
How AllthingsPM does this. The free Foundations chapter opens with context windows and attention and pretraining, post-training and reasoning models, and explains why you cannot promise determinism. Each term links into the knowledge graph, so you see what it depends on.
Which prompting and context terms matter most?
| Term | What it means for a PM |
|---|---|
| Prompt | The input you send the model. |
| System prompt | Standing instructions that shape every response. |
| User prompt | The part of the input written by the end user. |
| Prompt engineering | Writing and testing prompts to get reliable behavior. |
| Context engineering | Deciding what information goes into the window, and in what order. |
| Zero-shot prompting | Asking for a task with no examples. |
| Few-shot prompting | Including a few worked examples in the prompt [3]. |
| Chain-of-thought | Asking the model to reason step by step [3]. |
| Prompt template | A reusable prompt with slots for variables. |
| Prompt chaining | Splitting a task into several prompts run in sequence. |
| Structured output | Forcing replies into a schema such as JSON. |
| JSON mode | A setting that guarantees parseable JSON. |
| Stop sequence | Text that tells the model to stop generating. |
| Max tokens | The cap on output length, and on output cost. |
| Prefill | Starting the model's reply for it to steer format. |
| Persona | A role the prompt asks the model to adopt. |
| Grounding | Tying answers to supplied sources rather than memory. |
| Prompt caching | Reusing a repeated prompt prefix to cut cost and latency. |
| Context rot | Quality dropping as the window fills with noise. |
| Compaction | Summarizing old context so a long task fits the window. |
How AllthingsPM does this. The lesson Context as a budget treats the system prompt as a product surface and shows when retrieval, not a longer prompt, is the fix. You then practice explaining that trade-off in a mock interview.
What do RAG and retrieval terms mean?
| Term | What it means for a PM |
|---|---|
| Retrieval-augmented generation (RAG) | Fetching relevant documents at runtime and passing them to the model [1]. |
| Knowledge base | The document set RAG searches. |
| Embedding | A vector that captures meaning, so similar items sit close together [3]. |
| Vector | A list of numbers representing text or images. |
| Vector database | A store built for fast similarity search over vectors. |
| Semantic search | Search by meaning rather than exact words. |
| Keyword search (BM25) | Classic word-match ranking. |
| Hybrid search | Combining semantic and keyword search. |
| Chunking | Splitting documents into retrievable pieces. |
| Chunk size | How big each piece is; too big dilutes, too small loses context. |
| Top-k | How many results retrieval returns. |
| Reranker | A second model that reorders retrieved results. |
| Retrieval recall | Share of relevant documents actually retrieved. |
| Retrieval precision | Share of retrieved documents that are relevant. |
| Citation | Showing which source supports each claim. |
| Metadata filtering | Narrowing retrieval by date, owner or permissions. |
| Index freshness | How current the searchable data is. |
| Agentic search | A model deciding what to search for, iteratively. |
| Deep research | Multi-step agentic search that produces a report. |
| Knowledge graph | Entities and their relationships stored as a network. |
How AllthingsPM does this. Retrieval intuition and the agentic search and Deep Research patterns are taught in chapter 6, with cost trade-offs spelled out. The AllthingsPM knowledge graph is itself a working example of the last term in this table.
Which AI agent terms come up in interviews?
| Term | What it means for a PM |
|---|---|
| AI agent | A model that directs its own steps and tool use toward a goal [4]. |
| Workflow | Model calls orchestrated through predefined code paths [4]. |
| Agentic | Describes systems with some autonomy over their next step. |
| Tool use (function calling) | The model requesting an external action, such as an API call. |
| Tool schema | The description that tells the model what a tool does and needs. |
| Harness | The loop, retries and stop conditions wrapped around the model. |
| Agent loop | Think, act, observe, repeat until done. |
| Stop condition | The rule that ends the loop. |
| Planning | The agent breaking a goal into steps. |
| Memory | What the agent keeps between steps or sessions. |
| Orchestrator | The component that routes work between models or agents. |
| Subagent | A helper agent given one slice of the task. |
| Multi-agent system | Several agents collaborating on one goal. |
| Model Context Protocol (MCP) | An open standard for connecting AI apps to data, tools and workflows [2]. |
| MCP server | A program that exposes tools or data over MCP [2]. |
| MCP client | The app side that connects to MCP servers [2]. |
| Computer use | An agent operating a screen, mouse and keyboard. |
| Human in the loop | A person approves or corrects agent actions. |
| Autonomy level | How much the agent may do without asking. |
| Task completion rate | Share of agent runs that finish the job correctly. |
How AllthingsPM does this. Chapter 6, Agents and agentic architecture, covers what is inside the harness and when multi-agent is worth it. For a shorter read, see agents vs workflows and the anatomy of an AI agent.
What training and customization terms should a PM understand?
| Term | What it means for a PM |
|---|---|
| Pretraining | Initial training on a large unlabeled corpus to predict the next word [1]. |
| Post-training | Everything after pretraining that makes a model useful and safe. |
| Fine-tuning | Further training a pretrained model on additional data [1]. |
| Supervised fine-tuning (SFT) | Fine-tuning on example inputs and ideal outputs. |
| RLHF | Training from human rankings of outputs [1]. |
| Reward model | A model that scores outputs to guide training. |
| Constitutional AI | Training against written principles rather than only human labels. |
| Distillation | A smaller model learning from a larger one [3]. |
| Quantization | Lowering weight precision to save memory and speed up inference [3]. |
| LoRA | A cheap fine-tuning method that trains small add-on weights. |
| Open-weight model | A model whose weights you can download and host. |
| Closed model | A model available only through the vendor's API. |
| Training data | The examples the model learns from. |
| Synthetic data | Training or test data generated by a model. |
| Data labeling | Humans marking the correct answer on examples. |
| Overfitting | Learning training data so well that new data suffers [3]. |
| Mixture of experts | Specialized sub-networks with a gate choosing which runs [3]. |
| Model card | A document describing a model's uses, limits and evaluations. |
| Data flywheel | Production usage producing data that improves the next model. |
| Knowledge cutoff | The date after which the model has no training data. |
How AllthingsPM does this. The course teaches the fix hierarchy first: prompt, context and harness before touching weights, a point made in the harness lesson. The data flywheel lesson shows how production turns into the next model.
Which eval and quality terms do AI PMs use daily?
| Term | What it means for a PM |
|---|---|
| Eval | A repeatable test of whether the AI output is good enough. |
| Golden set | A curated set of inputs with known good outputs. |
| Ground truth | The correct answer used for scoring [3]. |
| Offline eval | Testing before release on a fixed dataset. |
| Online eval | Measuring quality on live traffic. |
| LLM as a judge | Using a model to grade another model's output. |
| Rubric | Written criteria the grader scores against. |
| Human eval | People grading outputs, the slowest and most trusted signal. |
| Pairwise comparison | Picking the better of two outputs. |
| Benchmark | A public standard test used to compare models. |
| Benchmark saturation | Top models scoring so high a test stops separating them. |
| Precision | Correct positive predictions over all positive predictions [3]. |
| Recall | Share of actual positives the model catches [3]. |
| F1 score | The balance of precision and recall. |
| Hallucination | A confident answer that is false. |
| Faithfulness | Whether the answer matches the supplied sources. |
| Regression test | Checking a change did not break what worked. |
| Error analysis | Reading failures and sorting them into causes. |
| Failure attribution | Deciding which layer failed: model, context, harness or surface. |
| Evaluation awareness | A model behaving differently when it senses a test. |
How AllthingsPM does this. Chapter 8, Evals: define good and make the number defensible, is one of the deepest in the course, and the failure attribution lesson teaches the four-layer diagnosis. Pair it with the AI evals guide for PMs and evaluation awareness.
What cost, latency and infrastructure terms matter?
| Term | What it means for a PM |
|---|---|
| Latency | Time between sending a prompt and getting the reply [1]. |
| Time to first token (TTFT) | How fast the first token appears [1]. |
| Tokens per second | Output speed once generation starts. |
| Streaming | Showing output as it is generated. |
| Input tokens | Tokens you send; billed per token. |
| Output tokens | Tokens the model writes; usually priced higher. |
| Cost per task | Total model spend to complete one user job. |
| Batch API | Asynchronous processing at a discount for non-urgent work. |
| Rate limit | Cap on requests or tokens per minute. |
| Throughput | How much work the system handles at once. |
| GPU | The chip most models train and run on. |
| Model routing | Sending each request to the cheapest model that can handle it. |
| Fallback model | A backup model used when the first fails. |
| Caching | Storing repeated results to save cost and time. |
| Edge inference | Running a model on the device. |
| Self-hosting | Running an open-weight model on your own servers. |
| API | The interface your product uses to call the model. |
| SDK | A code library that wraps the API. |
| Observability | Logging and tracing each model call. |
| Trace | The full record of one request, including tool calls. |
How AllthingsPM does this. Chapter 2, Data fluency, teaches you to read logs and traces yourself, and chapter 9 covers unit economics and pricing. These are the numbers interviewers push on in AI product sense rounds.
What safety, security and governance terms should an AI PM know?
| Term | What it means for a PM |
|---|---|
| Guardrail | A check that blocks or fixes unsafe input or output. |
| Alignment | Making model behavior match intended goals and values. |
| HHH | Helpful, honest, harmless: Anthropic's research framework [1]. |
| Red teaming | Deliberately trying to make the system fail. |
| Jailbreak | A prompt that bypasses safety behavior. |
| Prompt injection | Input that alters the model's instructions; OWASP's top LLM risk [5]. |
| Indirect prompt injection | Malicious instructions hidden in retrieved content. |
| Sensitive information disclosure | The model leaking private data [5]. |
| System prompt leakage | Users extracting hidden instructions [5]. |
| Excessive agency | Giving an agent more power than the task needs [5]. |
| Data poisoning | Corrupting training, fine-tuning or embedding data [5]. |
| Unbounded consumption | Runaway usage that drives cost or outages [5]. |
| Content moderation | Filtering harmful content. |
| Bias | Systematic unfair differences in outputs across groups. |
| Explainability | Being able to say why the model produced an output. |
| PII | Personally identifiable information. |
| Data retention | How long prompts and outputs are stored. |
| Audit log | A tamper-evident record of who did what. |
| Responsible AI | Practices for building AI that is fair, safe and accountable. |
| AI governance | Policies and owners for how AI is built and used. |
How AllthingsPM does this. The AI PRD lesson makes you name risks, guardrails and success metrics before anything is built. Real questions such as how Claude Code should handle hundreds of parallel subagents test this area directly.
Which AI UX and product terms show up in specs?
| Term | What it means for a PM |
|---|---|
| Copilot | AI that assists while the human stays in charge. |
| Autopilot | AI that completes the task on its own. |
| Chat interface | A conversational UI. |
| Ambient AI | AI that acts in the background without being asked. |
| Suggested prompts | Starter prompts that teach users what is possible. |
| Confidence indicator | UI that signals how sure the system is. |
| Graceful failure | A useful response when the AI cannot help. |
| Escalation | Handing a case to a human. |
| Feedback loop | Thumbs, edits and corrections captured as data. |
| Undo | Letting users reverse an AI action. |
| Approval step | A human check before an agent acts. |
| Trust calibration | Users trusting the AI as much as it deserves, no more. |
| Overreliance | Users accepting wrong AI output without checking. |
| Personalization | Adapting output to a specific user. |
| AI-native product | A product that could not exist without the model. |
| Wrapper | A thin product on top of someone else's model. |
| Moat | What stops rivals copying you, such as data or distribution. |
| Model dependency | Risk from relying on one vendor's model. |
| Build vs buy | Choosing to train, fine-tune or call an API. |
| Prototype | A quick build to test whether the AI can do the job. |
How AllthingsPM does this. Chapter 7, AI UX and human oversight, is about designing for a system that is wrong sometimes. The problem-first test lesson asks whether an agent is even the right answer.
What metrics, business and career terms complete the list?
| Term | What it means for a PM |
|---|---|
| AI product manager | A PM who owns products whose core behavior comes from a model. |
| Agent product manager | A PM who owns an AI agent's outcomes, often in customer service. |
| Forward deployed engineer | An engineer embedded with customers to ship AI into their systems. |
| AI PRD | A spec that defines quality bars, risks and evals, not just features. |
| Success metric | The number that says the feature worked. |
| Guardrail metric | A number that must not get worse. |
| Resolution rate | Share of issues the AI solves without a human. |
| Deflection rate | Share of contacts that never reach a human. |
| Acceptance rate | Share of AI suggestions users keep. |
| A/B test | Comparing two versions on live users. |
| Adoption | Share of eligible users who use the feature. |
| Retention | Share of users who come back. |
| Gross margin | Revenue left after model and serving costs. |
| Seat pricing | Charging per user. |
| Usage pricing | Charging per token, call or task. |
| Outcome pricing | Charging per result, such as a resolved ticket. |
| Brownfield deployment | Shipping AI into an existing company's messy systems. |
| Pilot | A limited rollout to prove value. |
| Product sense | Judgment about what to build and why. |
| AI PM interview loop | The rounds AI PM candidates face, from product sense to technical. |
How AllthingsPM does this. Chapter 9 teaches seat, usage and outcome pricing, chapter 10 covers brownfield enterprise deployment, and chapter 14, Get the job, walks the AI PM interview loop. Then practice on one of 116 live AI company job descriptions in the jobs catalog, such as the Product Lead, AI/ML (Evals) role at Abridge, or paste any JD into the JD mock interview.
How should you actually learn these 200 terms?
Learn them in the order a real product forces on you, which is the order of the ten tables above. Then test yourself out loud. An interviewer will not ask you to define RAG. They will ask why your support bot gives wrong refund answers, and expect chunking, retrieval recall, grounding and an eval in the same breath.
How AllthingsPM does this. Every course chapter ends in a graded case study. The question bank has 4,122 real questions from 260 companies, each with an answer guide, and any one can start a scored mock.
Why AllthingsPM is the better choice for learning AI PM vocabulary
Free glossaries from The Product Space (120 terms), AI Product Craft (100) and Product Leadership (80 plus) are clean references worth a bookmark [6][7][8]. But a list cannot tell you when your answer is wrong. AllthingsPM adds three things a glossary cannot:
- A course built from demand. 14 chapters and 101 lessons, derived from 604 real PM job postings and updated weekly, so the terms you learn are the ones employers ask for.
- A map of how ideas connect. The knowledge graph links 212 AI concepts through 1,113 connections, so you see that evals depend on ground truth, which depends on labeling.
- Practice that scores you. Mock interviews built from any job description, in text or voice, with follow-ups and a score, plus 4,122 real questions.
All of it sits in one account with a free tier (one JD mock a day) and Pro at $20 a month or $120 a year. For anyone serious about getting into AI product management, the verdict is simple: bookmark a glossary, but learn on AllthingsPM. Start the AI PM course free.
Frequently asked questions
What is the best AI product management glossary?
AllthingsPM's glossary is the best starting point because each of its 200 terms links to a lesson in an AI PM course built from 604 real job postings, and its knowledge graph maps 212 concepts. The Product Space and AI Product Craft offer useful free lists of 120 and 100 terms.
What AI terms should a product manager learn first?
Start with LLM, token, context window, RAG, embeddings, hallucination, evals, agents, guardrails and inference cost. These map directly to decisions about quality, cost and safety. The free Foundations chapter covers the first group.
What is the difference between an AI agent and a workflow?
Anthropic defines workflows as model calls orchestrated through predefined code paths, and agents as systems where the model dynamically directs its own process and tool use [4]. Workflows are more predictable; agents handle open-ended tasks.
Do AI PMs need to know how to code?
Most AI PM roles do not require production coding, but they do expect you to read logs, run basic SQL and reason about evals. AllthingsPM's Data fluency chapter teaches exactly that level.
What is an eval in AI product management?
An eval is a repeatable test of whether a model's output is good enough for your product, scored against a golden set, a rubric or human judgment. It is how AI PMs turn "it seems better" into a defensible number.
How do I use these terms in an interview?
Use them inside a diagnosis, not as definitions. Practice with a JD mock interview built from the role you want, and check your resume uses the right vocabulary with a resume review against the JD.
Ready to use the vocabulary?
Knowing 200 words is not the job. Making a trade-off with them is. Start the AllthingsPM AI PM course free and take your first scored mock the same day.
Sources
- Anthropic, Glossary, Claude documentation, accessed 29 September 2026.
- Model Context Protocol, What is the Model Context Protocol (MCP)?, accessed 29 September 2026.
- Google for Developers, Machine Learning Glossary, accessed 29 September 2026.
- Anthropic, Building effective agents, December 2024.
- OWASP GenAI Security Project, 2025 Top 10 for LLM Applications, accessed 29 September 2026.
- The Product Space, 120 AI Terms Every Product Manager Should Know in 2026.
- AI Product Craft, The Top 100 AI and ML Terms AI Product Managers Need to Know.
- Product Leadership, AI Product Management Glossary: 80+ Key Terms.
- AI Product Management Coaching, The Ultimate AI Product Management Glossary: 20+ Terms, Medium.
- AllthingsPM platform counts (chapters, lessons, concepts, connections, questions, job descriptions), September 2026.




