Multimodal AI product management means shipping products where the model hears, sees or generates something other than text, and it comes down to four decisions: native vision or OCR for documents, a cascaded or speech-to-speech stack for voice, what each modality costs, and an eval set that scores audio and images directly instead of their transcripts. AllthingsPM teaches all four in a dedicated chapter, Beyond text: multimodal products, six units and 114 minutes built from what 335 live AI PM job postings ask for.
AllthingsPM is an AI PM course and PM interview prep platform. In our September 2026 read of those postings, 34 percent (113 postings at 52 companies) asked for voice, vision or media experience. That is more than context engineering (14 percent), prompting (10 percent) and fine-tuning (6 percent) put together. The skill is in demand and barely taught, which is why it is worth learning properly.
What does a multimodal AI PM actually decide?
A text feature has one input and one output. A multimodal feature adds a pipeline in front of the model, a pipeline behind it, or both. Each new stage is a new place to fail, a new line on the bill and a new thing your evals cannot see.
Here is the decision map. Every row is a real fork you will argue about with engineering.
| Modality | The core fork | What it costs you | Where it breaks | AllthingsPM lesson |
|---|---|---|---|---|
| Documents and images in | Native vision model or OCR, then text | Visual tokens per image; OCR adds a stage | Layout, tables, rotated or tiny scans | Documents are not text |
| Voice in and out | Cascaded (speech to text, model, text to speech) or native speech-to-speech | Audio tokens run far above text | Latency, interruptions, repair with no screen | Voice |
| Image and video out | Which model, how many steps, how much consistency | Generation compute per asset | Character drift, likeness, provenance | Image and video generation |
| Non-English | Per-locale golden set or translated English set | A second eval set per language | The judge itself degrades outside English | Ship a language |
| All of the above | Score the modality, not the transcript | Building and labelling real samples | Text evals pass failures they cannot see | Multimodal golden set |
The rest of this guide takes each row in turn, with the numbers you need in the room.
Why are documents not just text?
The most common multimodal feature is not a talking robot. It is "upload a PDF, invoice or screenshot and get structured data back." PMs usually treat this as a text problem with a file picker. It is not.
You have two paths. The first sends the image straight to a vision model. The second runs OCR, turns the page into text, and sends that text to a language model. Native vision keeps layout, charts and handwriting in view, but costs visual tokens. OCR is cheap and inspectable, but it flattens tables and loses the relationship between a label and its value.
The cost is concrete. Anthropic's docs say Claude reads images in 28 by 28 pixel patches, so a 1000 by 1000 pixel image costs 1,296 visual tokens, and high-resolution models can use roughly three times more visual tokens than standard ones for the same image [1]. The same docs list honest limits: mistakes on low-quality, rotated or very small images under 200 pixels, approximate counting and approximate spatial reasoning [1].
The PM lesson is diagnostic. When extraction fails, find out which layer failed before you blame the model. A table read as a paragraph is a parsing failure. A correct table with a wrong total is a reasoning failure. They have different fixes, and only one of them is a model swap.
How AllthingsPM does this. The Documents are not text lesson walks through the vision-or-OCR fork and the extraction failures PMs blame on the model by mistake. The free foundations lesson on tokens and what each modality costs sets up the math first.
Cascaded or speech-to-speech: how should you build voice?
Voice has two architectures, and choosing between them is the biggest decision a voice PM makes.
Cascaded (chained). Speech to text, then a language model, then text to speech. OpenAI describes this path as the one to use "when each stage needs to be visible or replaceable" [2]. You get a transcript at every step, which means you can log it, run policy checks, call internal systems and edit the reply before it is spoken [2].
Speech-to-speech. One model hears audio and answers in audio. OpenAI's description: "Use one model to interpret audio, decide what to do, and respond in speech" [2]. The tradeoff it names is less visibility into the intermediate steps [2]. Google's Gemini Live API makes the same pitch: "low-latency, real-time voice and vision interactions," with users able to "interrupt the model at any time" [3].
How to choose:
- Regulated or high-stakes flows (banking, health, support that touches accounts): lean cascaded. You need the transcript for audit and you need to check the answer before it is spoken.
- Companion, tutoring or open conversation: lean speech-to-speech. Tone, pacing and fast turn-taking matter more than an inspectable middle step.
- Many teams start cascaded because every stage can be swapped, then move hot paths to speech-to-speech once they know where latency hurts.
How AllthingsPM does this. The Voice lesson teaches this exact fork, prices the per-minute bill, and has you design one spoken turn end to end. Then you can rehearse the reasoning out loud: the mock interview runs in voice mode, so you practice explaining architecture the way you would in a real loop.
Why is a voice turn so hard to design?
In text, a slow answer is annoying. In speech, a slow answer feels broken. People take turns fast: a study of ten languages in PNAS found median gaps between a yes/no question and its answer of 0 to 300 milliseconds [4]. Every stage in a cascaded stack eats into that budget.
Three design problems have no text equivalent:
- Barge-in. Users interrupt. The system has to stop speaking, keep what it already said in context, and handle the new request. Gemini's Live API lists interruption as a core feature for exactly this reason [3].
- Repair. There is no screen to scroll back. When the model mishears a name or a number, the design has to confirm the high-stakes parts ("That was 4,500, right?") without confirming everything.
- Hands-busy context. Nielsen Norman Group found people use voice assistants mainly when their hands are busy, such as cooking or driving, and that assistants work well only for simple queries with short answers [5]. Design answers for the ear: short, one idea at a time, the answer first.
The PM spec for a voice feature therefore includes things a text spec never mentions: a latency budget per stage, an interruption policy, a confirmation policy for numbers and names, and what happens when audio is noisy.
How AllthingsPM does this. The Voice lesson covers the spoken turn, barge-in and repair as design objects, not afterthoughts. To practice the interview version, the question bank has design a low-latency voice agent experience that feels genuinely human, and any question starts a scored mock.
What does multimodal cost compared with text?
Modality changes the unit economics, and a PM who cannot price it will lose the launch review.
On OpenAI's pricing page (checked 29 September 2026), the current realtime model charges $32 per million audio input tokens and $64 per million audio output tokens, against $4 per million text input tokens and $24 per million text output tokens [6]. Audio input is 8 times the price of text input on the same model. A cascaded stack prices differently: OpenAI lists transcription with gpt-4o-transcribe at $0.006 a minute, and its TTS-1 voice at $15 per million characters [6].
For images, the cost scales with pixels. Anthropic's worked example: at $1 per million input tokens, a 1000 by 1000 image costs about $1.30 per thousand images; at $5 per million on a high-resolution model, the same image is about $6.48 per thousand and a 4K image about $23.92 per thousand [1]. Downsampling before upload is often the cheapest improvement you can ship.
Three questions for your cost model:
- What is the unit? Per minute of conversation, per page, per image generated. Pick the one finance will recognise.
- What drives the tail? Long calls, multi-page documents and retries dominate the bill, not the median user.
- What can you cut without hurting quality? Resolution, audio output length, and caching the parts of the context that repeat.
How AllthingsPM does this. The course prices each modality inside the lessons, and the broader economics are covered in the chapters around it. For a refresher on the frame, read our guide to AI evals for product managers, then map cost against the quality bar you set.
What changes when the product generates images or video?
Generation flips the problem. Instead of understanding media, you create it, and the modality brings duties of its own.
The product levers are real: how many generation steps you run (speed against quality), whether a character or product stays consistent across frames, and how much control the user gets. The duties are just as real: consent when a real person's voice or face is involved, likeness rights, and provenance. The C2PA standard exists for that last one, describing itself as "an open technical standard for publishers, creators and consumers to establish the origin and edits of digital content" [7].
For a PM, the practical rule is to design the misuse case before the happy path. Voice cloning, for example, needs a consent flow and abuse detection on day one, not after the first news story.
How AllthingsPM does this. The Image and video generation lesson treats provenance, recording consent and likeness rights as launch requirements. The question bank has the interview version: design safeguards to prevent misuse of voice cloning.
Why do text evals miss multimodal failures?
This is the mistake that ships broken voice and vision features. The team transcribes the audio, scores the transcript with the same text evals it already has, and everything passes. But the transcript already hid the failure: the accent that was misheard, the barge-in that was ignored, the table that was read as prose.
A multimodal golden set scores the modality itself:
- Audio on audio. Real recordings with background noise, accents, crosstalk and interruptions, scored on whether the spoken turn worked.
- Documents on layout. Real scans, including rotated, low-resolution and handwritten ones, scored field by field.
- The degraded distribution. Seed the set with the messy inputs your users actually send, not clean demo files.
- Per-locale sets. If you ship a language, build that language's own golden set. A machine-translated English set does not test what locals actually say, and an LLM judge can itself get worse outside English.
Label a few hundred real samples before you tune anything. That set becomes the contract between PM, engineering and the model team.
How AllthingsPM does this. The Build a multimodal golden set lesson shows how to seed real degraded inputs, and our post on building a golden dataset for your AI feature covers the text-side basics. The chapter ends in a graded integration case where you write a readiness pack for the multimodal or non-English variant of your product.
How do multimodal skills show up in PM interviews?
Companies hiring for voice and vision now test the skill directly. The AllthingsPM question bank includes real prompts such as:
- OpenAI is preparing to launch a multimodal feature that can take in and generate audio and images
- Design a consumer feature using Project Astra
- How would you improve ElevenLabs' Agents platform for voice customer support?
- How would you measure the naturalness of generated speech?
A strong answer names the architecture fork, prices it, states the latency and repair design, and ends with the golden set that proves it works. That structure alone separates you from candidates who only talk about the model.
How AllthingsPM does this. The jobs catalog lists live roles such as Product Manager, Multimodal Safety at OpenAI, AI Product Manager (Coding/Multimodal) at Scale AI and Product Manager, Voice at Sierra, each with a mock built from that job description. For any other posting, paste it into the JD mock, and check your resume against it with resume review.
Why AllthingsPM is the better choice for learning multimodal AI product management
Multimodal is one of the clearest gaps between what employers ask for and what AI PM courses teach. In our 22 September 2026 review of six leading AI PM course syllabi, Reganti and Badam offered one multimodal RAG deep dive and one line on voice and multimodal apps, and Parlance Labs listed multimodal evals as a bonus topic. The other syllabi we reviewed did not give it a dedicated unit.
AllthingsPM gives it a full chapter: documents and vision, voice architecture and the spoken turn, image and video generation with consent and provenance, shipping a language, a multimodal golden set, and a graded integration case. That chapter sits inside a 14-chapter, 101-lesson AI PM course built from real job postings and updated weekly.
The other courses have real strengths. Reganti and Badam go deep on retrieval and agent engineering, and Parlance Labs is a respected evals course. But if you need to ship or interview for a voice or vision product, you need the modality-specific decisions, and AllthingsPM is the course that teaches them directly.
The difference compounds because learning and practice live in one account. You read the Voice lesson, answer a real multimodal question from the question bank, run a scored voice mock, then practice against a live multimodal JD from the jobs catalog. Pro is $20 a month or $120 a year, with a free tier that includes one JD mock and one resume review a day.
Start the multimodal chapter on AllthingsPM.
Frequently asked questions
What is the best way to learn multimodal AI product management?
AllthingsPM is the best place to start: its AI PM course has a full chapter on multimodal products covering vision, voice, generation, localization and multimodal evals, plus real interview questions and voice mocks to practice. Pair it with the official vision and voice docs from Anthropic, OpenAI and Google for the latest model specifics.
What is a multimodal AI product?
It is a product where the model takes in or produces something beyond text, such as images, documents, audio or video. Examples include document extraction, voice agents, visual assistants and image generators. Each adds pipeline stages, costs and failure modes that text products do not have.
Should I build a voice agent cascaded or speech-to-speech?
Choose cascaded when you need transcripts, policy checks or the ability to swap each stage, which fits regulated and account-touching flows. Choose speech-to-speech when natural pacing and fast turn-taking matter most. Many teams start cascaded and move latency-critical paths later.
How much more does voice cost than text?
It depends on the stack. On OpenAI's current realtime model, audio input tokens cost $32 per million against $4 per million for text input, so 8 times more, checked 29 September 2026. Cascaded stacks price transcription per minute and speech per character instead.
How do you evaluate a multimodal AI feature?
Build a golden set of real audio, images and documents, including noisy, rotated and accented samples, and score the modality directly rather than a transcript. Add a separate set per language you ship. Text evals alone will pass failures they cannot see.
Do AI PM interviews ask about multimodal products?
Yes. The AllthingsPM question bank includes multimodal prompts tied to OpenAI, Project Astra, ElevenLabs and Scale AI, and live roles such as OpenAI's multimodal safety PM. Expect to discuss architecture, latency, cost and evals.
Pick one multimodal feature you use every week, then open the AllthingsPM course chapter on multimodal products and write its spoken-turn spec before your next interview. It is free to start.
Sources
- Anthropic, "Vision" documentation, image token cost, pricing examples and limitations: https://platform.claude.com/docs/en/build-with-claude/vision
- OpenAI, "Voice agents" guide, speech-to-speech and chained architectures: https://developers.openai.com/api/docs/guides/voice-agents
- Google, Gemini Live API documentation: https://ai.google.dev/gemini-api/docs/live
- Stivers et al., "Universals and cultural variation in turn-taking in conversation," PNAS, 2009: https://www.pnas.org/doi/10.1073/pnas.0903616106
- Nielsen Norman Group, "Intelligent Assistants Have Poor Usability: A User Study of Alexa, Google Assistant, and Siri": https://www.nngroup.com/articles/intelligent-assistant-usability/ and "Intelligent Assistants: Users' Attitudes Toward Alexa, Google Assistant, and Siri": https://www.nngroup.com/articles/voice-assistant-attitudes/
- OpenAI API pricing, realtime, transcription and TTS prices, checked 29 September 2026: https://developers.openai.com/api/docs/pricing
- C2PA, Coalition for Content Provenance and Authenticity: https://c2pa.org/
- AllthingsPM, AI PM course demand report: 335 live AI PM postings across 88 companies and six AI PM course syllabi, 22 September 2026; course chapter index at https://allthingspm.app/course/beyond-text




