Use a workflow when you can write down the steps before the model runs. Use an agent only when the next step depends on what the model discovers along the way. Workflows are cheaper, faster and easier to test; agents handle open-ended work at a higher cost in tokens, latency and risk. Most good AI products are workflows with one small agent inside. The quickest way to learn to make this call is AllthingsPM, whose AI PM course opens its agents chapter with a lesson on exactly this decision: if you can write the path, it is a workflow.
AllthingsPM is an AI PM course and PM interview prep platform. Its course was built from 604 real PM job postings, and the decision in this guide is one employers now test directly.
What is the difference between an AI agent and a workflow?
The cleanest definitions come from Anthropic's engineering team. Workflows are "systems where LLMs and tools are orchestrated through predefined code paths." Agents are "systems where LLMs dynamically direct their own processes and tool usage, maintaining control over how they accomplish tasks" [1].
In PM terms: in a workflow, your team decides the order of steps. The model fills in each step (summarize this, classify that, draft a reply), but it never picks what happens next. In an agent, the model runs in a loop. It looks at the goal, chooses a tool, reads the result, and decides what to do next, until it judges the task done or hits a limit.
OpenAI draws the same line from the other side. Its guide says apps that "integrate LLMs but don't use them to control workflow execution" are not agents, and names simple chatbots, single-turn LLMs and sentiment classifiers as examples [2]. An agent "leverages an LLM to manage workflow execution and make decisions" and "dynamically selects the appropriate tools depending on the workflow's current state" [2].
Here is the whole comparison on one page.
| Question | Workflow | Agent |
|---|---|---|
| Who decides the next step? | Your code, written in advance | The model, at run time |
| Can you draw the flowchart before it runs? | Yes | No, it depends on what the model finds |
| Cost per run | Roughly fixed and predictable | Varies; agents use about 4x the tokens of chat [3] |
| Latency | Known upper bound | Open-ended until a stop condition fires |
| How you test it | Check each step's output | Check the final state, over many repeated runs [5] |
| Typical failure | A step gives a weak output | The loop wanders, repeats, or takes a wrong action |
| Best for | Known, repeatable, low-ambiguity tasks | Open-ended tasks where the number of steps cannot be predicted [1] |
| Example | Summarize a ticket, tag it, route it to a queue | Investigate a bug across logs, code and docs, then propose a fix |
Definitions from Anthropic [1] and OpenAI [2]; token figure from Anthropic's multi-agent research post [3].
What is the one-question test for agents vs workflows?
Ask: can I write the path down before the model runs?
If yes, it is a workflow, even if the steps are long and the model does most of the work inside each step. Anthropic's first piece of advice is to "find the simplest solution possible, and only increase complexity when needed" [1]. A written path is simpler to build, to price and to debug.
If no, because the right next step depends on what the model reads in a log file, a customer's reply or a search result, you may need an agent. Anthropic reserves agents for "open-ended problems where it's difficult or impossible to predict the required number of steps" [1].
OpenAI gives three signals that a task is worth an agent [2]:
- Complex decision-making: "nuanced judgment, exceptions, or context-sensitive decisions", such as refund approval in customer service.
- Difficult-to-maintain rules: rulesets that have become "unwieldy", such as vendor security reviews.
- Heavy reliance on unstructured data: interpreting natural language or documents, such as processing a home insurance claim.
OpenAI's advice is to "validate that your use case can meet these criteria clearly" before committing to an agent [2]. If none of the three applies, you are paying agent prices for workflow work.
How AllthingsPM does this: the course lesson Workflow or agent: if you can write the path, it is a workflow walks you through this exact test on real product cases. It builds on a discovery lesson that asks the question even earlier, is an agent even the right answer?, so you learn to challenge the "let's build an agent" request before a line of code exists.
What are the five workflow shapes every PM should know?
Before reaching for an agent, check whether one of these five patterns from Anthropic already fits [1]. Each is a workflow because your code, not the model, decides the order.
| Shape | What it does | PM example |
|---|---|---|
| Prompt chaining | Splits a task into fixed steps, each model call using the last one's output | Draft a release note, then check it against a style guide, then translate it |
| Routing | Classifies the input and sends it to a specialized path | Send billing tickets, bug reports and refund requests to different handlers |
| Parallelization | Runs several model calls at once and combines them | Score one answer on accuracy, tone and safety at the same time |
| Orchestrator-workers | A central model breaks a task into subtasks and hands them out | Plan edits across many files, then assign each file to a worker call |
| Evaluator-optimizer | One call generates, another critiques, in a loop | Write a summary, grade it against a rubric, revise until it passes |
The last two sit close to the agent line. An orchestrator decides the subtasks at run time, which is why Anthropic notes it suits tasks where "you can't predict the subtasks needed" [1]. Reaching for it is a sign you are one step from an agent.
How AllthingsPM does this: the same lesson covers the five shapes, and the next one, the anatomy of an agent, shows what you add when you cross the line: tools, planning, state, termination and escalation. The AllthingsPM knowledge graph shows how these concepts connect to evals, tool design and trust across the course.
Why are PMs being asked to make this call now?
Because employers write it into job descriptions. We counted mentions in the 677 PM job postings from 103 companies in the AllthingsPM JD corpus. 246 of them (36%) name agents or agentic work, and 230 (34%) name workflows. Among AI-native postings, 199 of 335 (59%) name agents, against 47 of 342 (14%) for other PM roles [6].
Some companies now hire for the role by name. Decagon, for example, lists a Senior Agent Product Manager and a Product Manager, Enterprise Agent Platform.
There is also a cost to getting it wrong. In June 2025 Gartner predicted that over 40% of agentic AI projects will be canceled by the end of 2027, due to escalating costs, unclear business value or inadequate risk controls [4]. Gartner also warned about "agent washing", rebranding assistants, RPA and chatbots as agents, and estimated only about 130 of the thousands of agentic AI vendors are real [4]. A PM who can say "this is a workflow, and here is why" saves the team from being one of those projects.
How AllthingsPM does this: the course is built from these same postings, so the agents chapter exists because the demand is real, not because the topic is fashionable. Open any AI company role in the AllthingsPM jobs catalog to see the agent language in context, then start a mock interview built from that posting.
What does an agent really cost compared with a workflow?
Anthropic states the trade plainly: "Agentic systems often trade latency and cost for better task performance, and you should consider when this tradeoff makes sense" [1].
The numbers from its own research system are useful anchors. Agents "typically use about 4x more tokens than chat interactions", and "multi-agent systems use about 15x more tokens than chats" [3]. The payoff can be real: a multi-agent setup with Claude Opus 4 leading Claude Sonnet 4 subagents "outperformed single-agent Claude Opus 4 by 90.2%" on Anthropic's internal research eval [3]. And token usage alone explained 80% of the performance variance on the BrowseComp evaluation [3].
So the PM question is not "is the agent better?" It is "is the gain worth four to fifteen times the tokens on this task, at this volume, for this user?" For a deep research report a user waits minutes for, maybe yes. For tagging ten thousand tickets a day, almost never.
OpenAI's guide reaches the same place for team design: "maximize a single agent's capabilities first," because more agents "can introduce additional complexity and overhead, so often a single agent with tools is sufficient" [2].
How AllthingsPM does this: the lesson when multi-agent is worth it works through this cost math, with the agentic search and Deep Research patterns as cases. You leave able to write the cost line in a spec instead of discovering it on the invoice.
How do you test an agent differently from a workflow?
A workflow is tested step by step: did the classifier pick the right queue, did the summary keep the key facts. An agent has to be judged on the end state and on consistency, because the same request can take a different path each time.
The tau-bench paper made this concrete. It found that even state-of-the-art function-calling agents such as GPT-4o succeeded on fewer than 50% of its tasks, and were inconsistent: pass^8 was below 25% in the retail domain [5]. Pass^k measures whether the agent succeeds on all k repeated tries of the same task [5].
For a PM, that changes the launch bar. You need a reliability number (how often does it finish correctly across repeated runs), a list of actions it may take without asking, and a defined handoff. OpenAI recommends escalating to a human when an agent exceeds failure thresholds, such as failing to understand intent after multiple attempts, and for high-risk actions like "canceling user orders, authorizing large refunds, or making payments" [2].
How AllthingsPM does this: the evals chapter has a lesson on evaluating an agent, not an answer, using final-state assertions and pass-k reliability. The trust chapter covers what an agent may do without asking, weighing reversibility against reliability, which is the memo your security and legal teams will want. Our guide to AI evals for product managers goes deeper on the eval side.
What does the hybrid pattern look like in a real product?
Most shipped products mix the two. The workflow is the spine; the agent is a narrow step inside it where the path genuinely branches.
Take customer support. A workflow receives the ticket, routes it by type (routing), and pulls the customer's account (a fixed tool call). Only the hard branch, say a disputed refund with a long history, goes to an agent that can read past tickets, check policy and propose a resolution. Any refund above a set amount goes to a human, following OpenAI's high-risk rule [2].
The practical sequence for a PM:
- Write the path. Mark every step you can specify in advance.
- Circle the steps where the next move depends on what the model finds. Those, and only those, are agent candidates.
- For each, check OpenAI's three signals: judgment, unwieldy rules, unstructured input [2].
- Give the agent a tool list, a step or budget cap, and a stop condition.
- Define the human handoff and the actions that always need approval.
- Measure final-state success over repeated runs before launch [5].
How AllthingsPM does this: the chapter ends with an integration case, the agent version of your carried feature, where you design the architecture and tool contracts for a feature you carry through the whole course. The lesson on writing the tool contract covers the step most teams skip: workflow-shaped tools with capped output and errors written as instructions.
How does agents vs workflows come up in PM interviews?
At AI companies it shows up in product sense, strategy and execution rounds, usually disguised as a product question. Real examples from the AllthingsPM question bank:
- Sierra: a new German enterprise customer wants to automate support with an AI agent
- After launching an enterprise AI agent, what primary success metric and guardrail metrics would you track?
- Anthropic: Claude Code shows a recurring failure mode across many agentic coding tasks, such as looping
A strong answer names the decision out loud: which steps are fixed (workflow), which branch (agent), what the agent may do alone, where a human steps in, and how you will measure reliability.
How AllthingsPM does this: every question above has its own page and answer guide, and any of them can start a scored AI mock interview in text or voice, with follow-ups. For a specific role, paste the posting into the JD mock and the interviewer builds the loop around it. Our list of AI PM interview questions gathers more.
Why AllthingsPM is the better choice for learning agents vs workflows
You can read the vendor guides for free, and you should: Anthropic's and OpenAI's posts [1][2] are the clearest primary sources on this topic. What they do not do is teach you to make the call as a PM, test you on it, or connect it to the jobs you are applying for.
AllthingsPM does all three in one account. The AI PM course was built from 604 real PM job postings and is updated weekly. Its agents chapter takes you from the workflow-or-agent test through agent anatomy, tool contracts, the harness, multi-agent cost and MCP, and ends with an integration case where you design an agent for your own feature. Lessons in other chapters cover how to evaluate agents and what they may do without asking.
Then it connects learning to getting hired. The same account holds 4,122 real interview questions from 260 companies, each with an answer guide; 116 live PM job descriptions at 18 AI companies, each with a mock interview built from it; and a resume review against a job description so your agent project shows up on the page recruiters read.
General PM courses and interview sites have their strengths, such as brand recognition, live cohorts and large peer communities. For learning this specific AI PM skill and proving it in an interview, at $20 a month or $120 a year with a free tier to start, AllthingsPM is the stronger choice. Start the AI PM course free.
Frequently asked questions
What is the difference between AI agents and workflows?
A workflow runs LLMs and tools through code paths you define in advance; an agent lets the model decide its own next step and which tools to use [1]. If you can draw the flowchart before it runs, it is a workflow.
When should a PM choose an agent over a workflow?
When the number of steps cannot be predicted and the task needs judgment, unwieldy rules or unstructured input, the three signals in OpenAI's guide [2]. Otherwise a workflow is cheaper, faster and easier to test.
Are agents more expensive than workflows?
Usually. Anthropic reports agents use about 4x the tokens of chat interactions and multi-agent systems about 15x [3]. That spend can buy better results on hard tasks, so weigh it per use case.
Is an agentic workflow the same as an agent?
Not quite. An agentic workflow is usually a fixed path with one or more agent steps inside it. That hybrid is what most shipped products use: predictable spine, flexible branch.
What is the best way to learn agents vs workflows as a PM?
AllthingsPM is the best place to start: its AI PM course has a full chapter on agents and agentic architecture, opening with the workflow-or-agent decision, plus agent eval and trust lessons and real interview questions to practice on. Pair it with Anthropic's "Building effective agents" post [1].
Do PM interviews ask about agents vs workflows?
Yes, at AI companies. 59% of AI-native postings in the AllthingsPM JD corpus name agents [6], and questions about automating enterprise support with an AI agent appear in the AllthingsPM question bank.
Ready to make the call with confidence? Start the AllthingsPM AI PM course free, then practice it in a JD mock interview.
Sources
- Anthropic, "Building effective agents", 19 December 2024. https://www.anthropic.com/engineering/building-effective-agents
- OpenAI, "A practical guide to building agents". https://cdn.openai.com/business-guides-and-resources/a-practical-guide-to-building-agents.pdf
- Anthropic, "How we built our multi-agent research system", 13 June 2025. https://www.anthropic.com/engineering/multi-agent-research-system
- Gartner, "Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027", 25 June 2025. https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027
- Yao et al., "tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains", arXiv, 2024. https://arxiv.org/abs/2406.12045
- AllthingsPM JD corpus: 677 PM job postings from 103 companies, read 22 September 2026. https://allthingspm.app/jobs
- Prompt Engineering Guide, "AI Workflows vs. AI Agents". https://www.promptingguide.ai/agents/ai-workflows-vs-ai-agents




