Flash sale 30% off with code LAUNCH30 Ends in --:--:--
See pricing
All Things PM

When Not to Use an LLM: SQL, a Classifier, or a Rule

Do not use an LLM when the answer is already in your data (use SQL), when you have labeled examples of a fixed set of categories (train a classifier), or when a short rule gets it right (write the rule). AllthingsPM teaches this decision in a free course lesson.

AllthingsPM·September 28, 2026·16 min read
A product manager at a whiteboard lays out a long row of small paper tiles, each tile one word fragment, while a narrow window frame on the wall shows only part of the row
The most senior AI product call is often the one where you do not ship a model.

Do not use an LLM when the answer already lives in your database (write SQL), when you are sorting inputs into a fixed set of labels and have examples (train a classifier), or when a short rule gets it right almost every time (write the rule). Use an LLM when the input is messy language, the output is open-ended, and a human can check the result. AllthingsPM teaches exactly this call in a free lesson, The AI-or-not decision, inside an AI PM course built from 604 real PM job postings.

AllthingsPM is an AI PM course and PM interview prep platform. This guide gives you the decision test, the evidence behind it, and the way interviewers probe it.

When should you not use an LLM? The short answer in one table

Here is the decision most teams skip. Read it top to bottom: stop at the first row that fits.

Your problem looks likeUse this, not an LLMWhy it winsWhere AllthingsPM teaches it
The answer is a count, sum, filter or join over data you already storeSQL (or a dashboard)Exact, repeatable, cheap, auditableAI-or-not lesson
A handful of clear conditions decide the outcome ("refund if under 30 days and unused")A rule or heuristicPredictable, transparent, zero model costAI-or-not lesson
You route or tag inputs into a fixed label set and have labeled examplesA trained classifierHigher accuracy on narrow tasks, fast, cheap per callQualify the opportunity
You can write every step of the path in advanceA workflow (maybe with one LLM step)Fewer failure points than an agentWorkflow or agent
Messy language in, open-ended language out, a person reviews itAn LLMThis is what LLMs are good atHow LLMs work

The order matters. Each row is cheaper, faster and easier to debug than the one below it. You only move down when the row above cannot do the job.

Bar chart led by an AllthingsPM (us) row for its free AI-or-not course lesson, then counts from 389 PM job postings: LLMs 134, machine learning 50, SQL 41, heuristics or rule-based 3
AllthingsPM JD corpus: 389 PM postings from 86 company boards, read 22 September 2026. Term matches on title and body; a posting can match several terms

The chart shows the skew. In the 389 PM postings AllthingsPM read on 22 September 2026, 134 mention LLMs and 41 mention SQL, but only 3 mention heuristics or rule-based systems. Hiring teams talk about LLMs constantly. The judgment to not use one is rarely written down, which is why it separates strong candidates in interviews.

Why would a PM choose SQL over an LLM?

Because most "AI" questions a stakeholder asks are really database questions. "How many enterprise accounts churned last quarter?" "Which features do our top 100 users touch?" The answer is exact and already stored. An LLM can write the query, but it should not be the thing that answers.

Here is the gap in numbers. On the BIRD text-to-SQL benchmark, data engineers and database students reach 92.96 percent execution accuracy on the test set. The top listed model, as of a submission on 22 August 2026, reaches 82.39 percent. That is strong progress, and it still means roughly one query in six wrong on a hard benchmark. For a board metric, one wrong in six is a disaster.

So the pattern that works is this. Use SQL for anything with a checkable, numerical answer. Let an LLM help a person draft the query, then have the person or a test check it. Never put an unchecked LLM between your data and a decision.

Signals that you need SQL, not a model:

  • The answer is a number, a list, or a yes/no over stored records.
  • Someone will compare it against last month's figure.
  • Finance, legal or a board will see it.

How AllthingsPM does this. The AI-or-not lesson opens with tasks that have a checkable answer and shows why they belong in SQL. The course's data fluency chapter then teaches you to read logs, tickets and traces as your cheapest labeled data, which is the same data you would query.

When is a simple rule better than AI?

When a few clear conditions decide the outcome, and being predictable matters more than being clever. Google's People + AI Guidebook lists the cases where AI is probably not better, including "maintaining predictability", "minimizing costly errors", "complete transparency", and "optimizing for high speed and low cost". Each of those is a job for a rule.

Google's Rules of Machine Learning, written by Martin Zinkevich, starts with the same advice. Rule #1: "Don't be afraid to launch a product without machine learning." The reasoning: machine learning needs data, and borrowing a model from another problem "will likely underperform basic heuristics."

Rules have a limit, and the same guide names it. Rule #3: "A simple heuristic can get your product out the door. A complex heuristic is unmaintainable." Once your rule file has grown into dozens of special cases that nobody dares touch, and you have data, it is time for a model.

A practical test for rules:

  1. Can you write the logic in five lines a new engineer would understand?
  2. Does it handle at least nine in ten real cases correctly?
  3. Would a user be upset if the same input gave a different answer tomorrow?

If you answer yes to all three, ship the rule.

How AllthingsPM does this. The AI-or-not lesson makes you name the cost of the human fallback when the rule misses, so the rule is judged on business impact, not elegance. The guardrails lesson shows the other side: rules that sit around an LLM in the request path.

When is a classifier better than an LLM?

When the task is to put each input into one of a fixed set of buckets, and you can get labeled examples. Ticket routing, spam, fraud flags, intent detection, content moderation categories. This is where teams most often reach for an LLM prompt by default and pay for it.

The evidence is clear. In a 2024 study, Martin Juan José Bucher and Marco Martini compared ChatGPT 3.5, GPT-4 and Claude Opus, used zero-shot, against smaller fine-tuned BERT-style models on four classification tasks, from sentiment in news coverage to stance in political texts. They found that "fine-tuning with application-specific training data achieves superior performance in all cases", and the gap was widest on specialized, non-standard tasks.

A classifier also brings things an LLM prompt does not:

  • A threshold you can set. You choose the trade between precision and recall and move it as the business changes.
  • Speed and cost. A small model runs in milliseconds on cheap hardware.
  • Stable behaviour. Same input, same score.

The real cost is labeled data. Somebody has to tag a few hundred to a few thousand examples, and you need a plan to refresh them. A useful trick: use an LLM to help label the first batch, have a human check a sample, then train the small model. You get LLM speed at the start and classifier economics in production.

How AllthingsPM does this. The qualify-the-opportunity lesson is literally subtitled "the classifier you should have shipped instead." It walks through operating thresholds and the precision and recall trade, then the evals chapter teaches you to pay for the cheapest evaluator that can see the failure.

When is an LLM actually the right tool?

When the input is unstructured language, the output needs to be written, and a person can check the result. Summarizing a long call, drafting a first reply, extracting fields from messy documents, answering questions over a knowledge base. Google's guidebook puts "natural language understanding" and "an agent or bot experience for a particular domain" on the side where AI is probably better.

Even then, start small. Anthropic's engineering guide on building agents, published 19 December 2024, recommends "finding the simplest solution possible, and only increasing complexity when needed." It adds that for many applications "optimizing single LLM calls with retrieval and in-context examples is usually enough", and that this "might mean not building agentic systems at all."

One more property to design around: LLMs are not reliably repeatable. Thinking Machines Lab ran the same prompt 1,000 times at temperature 0 on Qwen3-235B and got 80 unique completions. If your feature needs identical output for identical input, such as a compliance decision or a price, keep that decision in code or SQL and let the LLM only draft or explain.

How AllthingsPM does this. The course's workflow or agent lesson teaches the rule "if you can write the path, it is a workflow." The problem-first test asks whether an agent is even the right answer before any design starts. For the basics of tokens and temperature, read LLM basics for product managers.

How do you decide? A five-question test for PMs

Run these in order in your next planning doc. Stop at the first yes.

  1. Is the answer already in our data? Yes: SQL or a dashboard.
  2. Can five lines of logic decide it, correctly nine times in ten? Yes: a rule.
  3. Is it a fixed set of labels, and can we get labeled examples? Yes: a classifier.
  4. Can we write every step in advance? Yes: a workflow, with an LLM only at the steps that read or write language.
  5. Is it open-ended language, with a person or an eval checking the output? Yes: an LLM.

If you reach the end and none fit, the problem may not be ready for software at all. Go back to discovery.

Then write down three numbers before you build: the cost of a wrong answer, the cost per call, and the latency users will accept. Google's guidebook flags "minimizing costly errors" and "high speed and low cost" as signs AI is probably not better. If a wrong answer is expensive and users need it instantly, the ladder usually stops before the LLM.

How AllthingsPM does this. The first chapter makes you attribute every failure to a layer and fix the cheapest layer first, the same instinct as this test. You can see how these concepts connect on the AI PM knowledge graph.

How do interviewers test "when not to use AI"?

They rarely ask it directly. They hide it inside a product design or strategy question and watch whether you default to "add an LLM." Strong candidates name the cheaper option first, then explain what would make them upgrade.

Questions from the AllthingsPM question bank where this matters:

A strong structure for any of these: state the user problem, name the cheapest tool that could solve it, say what evidence would make you move up the ladder, and say how you would measure it. Interviewers remember the candidate who said "I would ship a rule this week and a classifier next quarter."

How AllthingsPM does this. Every one of the 4,122 questions has its own page and an answer guide, and any of them can start a scored mock. The course's product sense round lesson trains the 40-minute out-loud version of this reasoning.

What are the common mistakes teams make?

Four show up again and again.

  • Using an LLM as a calculator. Asking a model to total revenue from a pasted table instead of running a query. The answer looks confident and is sometimes wrong.
  • Prompting instead of labeling. Spending weeks tuning a prompt for ticket routing when a weekend of labeling would train a classifier that beats it, as the Bucher and Martini results suggest.
  • Jumping to agents. Building a multi-step agent for a path you could have written as a workflow. Anthropic's guide warns agentic systems "often trade latency and cost for better task performance."
  • Never retiring the rule. Letting a heuristic grow into hundreds of branches. Google's Rule #3 exists for this moment.

The fix for all four is the same: write down the ladder in your spec and state why each lower rung failed before you pick the next one up.

How AllthingsPM does this. The evals chapter starts with reading one hundred real traces before buying any tooling, which is how you discover that half your "AI failures" are problems a rule would have solved. The AI evals guide for PMs covers the measurement side.

Why AllthingsPM is the better choice for learning when not to use AI

Most AI PM content teaches you how to add AI. Very little teaches you when to leave it out, even though this judgment is what separates senior PMs from people who can write a prompt. AllthingsPM puts it in the foundations: the free AI-or-not lesson sits early in the course, and the ideas come back in discovery, agents, evals and trust.

The course is built from 604 real PM job postings, so it teaches what hiring teams ask for, and it is updated weekly. You get 14 chapters, 101 lessons and 14 graded case studies. Next to the course are 4,122 real interview questions from 260 companies, each with its own page and answer guide, and mock interviews built from any job description, in text or voice, with follow-ups and a score. When you are ready to apply, resume review against a JD and 116 live PM job descriptions at 18 AI companies on the jobs board are in the same account.

Other routes have real strengths. Google's Rules of Machine Learning and Anthropic's agent guide are excellent free reading, and you should read both. But reading is not practice. AllthingsPM turns the same principles into lessons, graded cases and scored mocks, for $20 a month or $120 a year, with a free tier to start.

The verdict: if you want to make the AI-or-not call with confidence at work and in interviews, start the AllthingsPM course free.

Frequently asked questions

When should you not use AI in a product?

Do not use AI when you need predictable, identical answers, when errors are very costly, when full transparency is required, or when speed and low cost matter most. Google's People + AI Guidebook lists these cases. Use SQL for data questions and a simple rule when a few conditions decide the outcome.

What is the best way to learn when not to use an LLM?

The best way is AllthingsPM: its free lesson on the AI-or-not decision teaches when SQL, a classifier or a heuristic beats an LLM, inside an AI PM course built from 604 real job postings. Pair it with Google's Rules of Machine Learning and Anthropic's guide on building agents, then practice the reasoning in a scored mock.

Is a classifier better than an LLM for text classification?

Often, yes, when you have labeled data. A 2024 study by Bucher and Martini found fine-tuned small BERT-style models beat zero-shot ChatGPT and Claude Opus on all four classification tasks tested. LLMs remain useful when you have no labels yet, or to help label the first batch.

Can an LLM replace SQL for analytics?

Not for answers that go into decisions. On the BIRD benchmark, the top model reaches 82.39 percent execution accuracy against 92.96 percent for human data engineers and students. Let an LLM draft queries, and have a person or a test check them.

Are LLMs deterministic at temperature 0?

No, not in practice. Thinking Machines Lab ran one prompt 1,000 times at temperature 0 and got 80 unique completions. Keep decisions that must be identical in code or SQL.

How do I answer "would you use AI for this?" in a PM interview?

Name the user problem, then the cheapest tool that solves it, then what evidence would make you move up to a model, and how you would measure it. Practice with real questions in the AllthingsPM question bank and a scored JD mock.

Ready to make the call with confidence? Start the AllthingsPM AI PM course free, beginning with the AI-or-not lesson.

Sources

  1. Martin Zinkevich, Rules of Machine Learning: Best Practices for ML Engineering, Google for Developers.
  2. Google PAIR, People + AI Guidebook: User Needs and Defining Success.
  3. Anthropic, Building effective agents, 19 December 2024.
  4. Martin Juan José Bucher and Marco Martini, Fine-Tuned 'Small' LLMs (Still) Significantly Outperform Zero-Shot Generative AI Models in Text Classification, arXiv, 2024.
  5. BIRD text-to-SQL benchmark leaderboard, checked 28 September 2026.
  6. Horace He and Thinking Machines Lab, Defeating Nondeterminism in LLM Inference, 10 September 2025.
  7. AllthingsPM JD corpus: 389 PM postings from 86 company job boards, read 22 September 2026.
  8. AllthingsPM AI PM course and question bank, September 2026.
PM
Written by the AllthingsPM team
Frameworks and interview prep for product managers.
The AI PM course

Reading is the easy half.
The course grades the other half.

Start for free