Flash sale 30% off with code LAUNCH30 Ends in --:--:--
See pricing
All Things PM

Make Your First LLM API Call as a PM (2026 Guide)

A product manager can make a first OpenAI API call in about fifteen minutes: get a key, set it as an environment variable, send one curl request to the Responses API, and read the usage block to learn what it cost. AllthingsPM teaches this in a free course lesson.

AllthingsPM·September 28, 2026·16 min read
A product manager at a desk types a short request into a terminal on a laptop while a paper envelope flies toward a distant server rack and a reply envelope returns
One request out, one answer back, and a usage block that tells you exactly what it cost.

You can make your first OpenAI API call as a product manager in about fifteen minutes, with no app and no framework: create an API key, store it as an environment variable, paste one curl command that posts to https://api.openai.com/v1/responses, and then read the usage numbers in the reply to see what the call cost. That single call teaches you more about tokens, latency and price than a week of reading.

AllthingsPM is an AI PM course and PM interview prep platform. Its free lesson Make the API call yourself walks you through exactly this exercise, inside a course built from 604 real PM job postings, so the skill connects straight to what hiring managers ask for.

How do you make your first OpenAI API call, step by step?

Here is the whole path. Each step is one small action, and none of them needs you to write a program.

StepWhat you doWhat it teaches you
0. Learn it in AllthingsPMOpen the free lesson Make the API call yourselfMessages, tokens, temperature, streaming and the usage block, in PM terms
1. Get a keyCreate an API key in your OpenAI accountKeys are billed credentials, not passwords you share
2. Store the keyexport OPENAI_API_KEY="..." in your terminalWhy keys live in the environment, never in code
3. Send one requestPaste the curl command belowThe request shape: model plus input
4. Read the responseFind the text and the usage numbersOutput is a list of items; tokens are the unit of cost
5. Change one thingAdd instructions, swap the modelHow behaviour and price move with each knob
6. Price a featureMultiply tokens by the price per millionUnit economics you can defend in a review

Step 1 and 2: get a key and keep it out of your code

Create a key in your OpenAI developer account and add a small amount of credit or a spend limit. Then set it in your terminal. OpenAI's quickstart gives this for macOS and Linux [1]:

export OPENAI_API_KEY="your_api_key_here"

On Windows PowerShell the quickstart uses setx OPENAI_API_KEY "your_api_key_here" [1]. OpenAI's SDKs read the key from the environment automatically [1].

Treat this key like a company card. OpenAI's own guidance is to keep keys server side, never put them in browser or mobile code, never commit them to a repository (even a private one), and revoke and replace a key you think has leaked [5]. The first time you paste a key into a shared doc "just for a second" is the time it leaks.

Step 3: send one request with curl

This is the exact first call from OpenAI's quickstart [1]:

curl "https://api.openai.com/v1/responses" \
    -H "Content-Type: application/json" \
    -H "Authorization: Bearer $OPENAI_API_KEY" \
    -d '{
        "model": "gpt-6-astra",
        "input": "Write a one-sentence bedtime story about a unicorn."
    }'

Read it like a PM, not an engineer. There are only four parts:

  • The endpoint (/v1/responses): the door you knock on.
  • Two headers: "I am sending JSON" and "here is my key."
  • The model: which brain you are paying for.
  • The input: what you are asking.

That is the whole contract. Every AI feature your team ships, however fancy, is some version of this request with more in the input.

Step 4: read the response, especially the usage block

The reply is JSON. OpenAI's text generation guide warns that the model's output lives in an output array that can hold several kinds of items, including tool calls and reasoning data, so "it is not safe to assume" the text sits at the first position [2]. The official SDKs add a convenience property, output_text, that joins all the text into one string [2].

Now find the usage object. It reports how many input tokens you sent and how many output tokens the model produced. That is your bill. Anthropic's Messages API returns the same idea with input_tokens and output_tokens, plus cache fields when prompt caching is on [4].

A token is a chunk of text. OpenAI's rule of thumb for English is that one token is about 4 characters or three quarters of a word, and 100 tokens is about 75 words [6]. Other languages can tokenize differently [6], which matters the day your product launches in a new market.

Bar chart from the AllthingsPM JD corpus: AllthingsPM read 389 AI company PM postings; 164 mention hands-on, 124 mention API or APIs, 87 prototype, 45 prompt, 42 evals, 41 SQL and 19 Python
Source: AllthingsPM JD corpus, 389 PM postings from 86 AI companies, read 22 September 2026

Why bother? Look at the chart. In the AllthingsPM JD corpus of 389 PM postings from 86 AI companies, 124 (32%) mention APIs and 164 (42%) ask for hands-on work, while only 19 (5%) mention Python. Hiring managers want PMs who can touch the product, not PMs who can write production code. One API call is the cheapest proof you can touch it.

How AllthingsPM does this. The free lesson Make the API call yourself sits in the Foundations chapter of the AllthingsPM AI PM course and covers messages, tokens, temperature, streaming and the usage block. It is followed by lessons on context windows and on attributing failures to a layer, so your first call becomes the base for everything after it.

What should you change after the first call works?

Change one thing at a time and watch what moves. This is product discovery on the model itself.

Add instructions. The Responses API accepts an instructions parameter that gives the model "high-level instructions on how it should behave," including tone, goals and examples [2]. Messages also carry roles: developer messages are prioritized ahead of user messages [2]. In product terms, instructions is your system prompt, the part the user never sees but that shapes every answer.

Try this version:

curl "https://api.openai.com/v1/responses" \
    -H "Content-Type: application/json" \
    -H "Authorization: Bearer $OPENAI_API_KEY" \
    -d '{
        "model": "gpt-6-luna",
        "instructions": "You are a support agent for a meal kit app. Answer in two sentences. If you are not sure, say so.",
        "input": "My box arrived warm. Is the chicken safe to eat?"
    }'

Run it five times. Do the answers differ? Does it ever say something your legal team would hate? You have just done your first informal eval, and you learned why "the model said something weird once" is not a bug report.

Swap the model. Change gpt-6-luna to gpt-6-sol, then to gpt-6-astra, and compare quality, speed and the usage numbers. The price gap between tiers is large, as the next section shows.

Try the same thing on Claude. Anthropic's Messages API posts to /v1/messages with an x-api-key header, and requires three fields: model, max_tokens and messages [4]. That max_tokens requirement is a useful lesson: you cap output length, and so output cost, on every call.

Know the temperature caveat. Temperature used to be the classic "creativity" knob. On Anthropic's API it is now deprecated: models released after Claude Opus 4.6 accept only the default of 1.0 and reject other values with a 400 error [4]. If a tutorial tells you to set temperature to 0 for consistency, check the docs for your model first.

How AllthingsPM does this. The course turns these experiments into skills you can show. Structured outputs explains why a JSON format is a decoding constraint rather than a polite request, and Context as a budget treats the system prompt as a product surface. When you are ready to measure quality properly, the Evals chapter picks up where your five manual runs stopped.

How much does an LLM API call cost?

Price is per million tokens, with separate rates for input and output. These are OpenAI's standard prices, checked 28 September 2026 on OpenAI's pricing page [3]:

ModelInput per 1M tokensOutput per 1M tokens
gpt-6-astra$10.00$50.00
gpt-6-sol$2.00$10.00
gpt-6-luna$0.10$0.50
gpt-5.4-mini$0.75$4.50
gpt-5.4-nano$0.20$1.25

Prices checked 28 September 2026 on OpenAI's own pricing page.

Work one example. Say your first call used 30 input tokens and 100 output tokens (your real numbers are in usage). On gpt-6-luna that costs 30 x $0.10 / 1,000,000 plus 100 x $0.50 / 1,000,000, which is about $0.00005. On gpt-6-astra the same call costs about $0.0053, roughly a hundred times more.

Now price a feature. Imagine a support assistant handling 10,000 conversations a day, each with 1,000 input tokens and 300 output tokens, on gpt-6-sol. Input is 10 million tokens a day, or $20. Output is 3 million tokens a day, or $30. That is $50 a day, about $1,500 a month, before retries, longer histories or tool calls. The same math is the core of estimation questions such as estimate Cursor's monthly LLM API cost per Pro user, which appears in the AllthingsPM question bank.

Two things surprise new AI PMs here. Output tokens usually cost several times more than input tokens, so long answers are the expensive part. And input grows silently: every turn of a chat resends the history, so turn ten costs more than turn one.

How AllthingsPM does this. The lesson Cost per successful task, the latency SLO, and the gross margin you defend teaches you to price by successful task rather than by call, and to route easy requests to cheaper models. You can then rehearse the numbers out loud in an AI mock interview and get scored follow-ups.

What errors will you hit, and what do they mean?

Your first call will probably fail once. That is normal, and the error codes are informative. From OpenAI's error code guide [7]:

  • 401, invalid authentication. The key is wrong, missing or revoked. Check it, or generate a new one.
  • 429, rate limit. You are sending requests too fast. Slow down and respect the Retry-After header.
  • 429, quota or spend limit. Same code, different cause: you are out of credit or hit your cap. Retrying will not help until billing changes.
  • 500, server error. The problem is on OpenAI's side. Wait, retry and check the status page.

Notice the product lesson inside the 429s. One code, two root causes, two completely different fixes. Every AI feature needs an error state for "slow down" and a different one for "we are out of budget," and a PM who has seen both will ask for both in the spec.

How AllthingsPM does this. The lesson Attribute every failure to a layer teaches you to decide whether a failure came from the model, the context, the harness or the surface. The Trust chapter then covers what happens when failures are deliberate, such as prompt injection.

Why does a PM need to make API calls at all?

Because AI products are built on this contract, and the decisions that matter are product decisions dressed as technical ones. Which model tier? How long can answers be? What goes in the instructions? What do we show when the provider returns a 429? You cannot weigh those trade-offs from a slide.

The job postings say so directly. In the AllthingsPM corpus, API mentions (124 postings) far outnumber Python mentions (19), and OpenAI itself lists roles such as Product Manager, API Agents and Product Manager, API Infrastructure, where the API is the product. For those roles, having made the call yourself is table stakes. For the rest, it is the difference between reviewing an AI spec and co-writing it.

It also changes how you talk with engineers. "The usage block shows 4,000 input tokens per turn, can we trim the history?" is a better conversation than "it feels slow and pricey." If you want the conceptual picture behind the call, read LLM basics for product managers; if you are wondering how far to take the coding, see do AI PMs need to code.

How AllthingsPM does this. Every AllthingsPM job page carries a mock interview built from that exact job description, so you can read the OpenAI API Agents posting and then practice being interviewed for it. The JD mock does the same for any posting you paste in.

What should you build after your first call?

Keep it small and useful. A few good second projects:

  1. A prompt comparison sheet. Run the same ten real user questions through two models and two instruction versions. Record answers, tokens and cost in a spreadsheet. This is the seed of an eval set.
  2. A structured extractor. Ask the model to turn messy feedback into JSON fields such as sentiment, feature and severity. You will quickly learn why format constraints matter.
  3. A cost model for one feature. Take your team's next AI idea and price it with real token counts from usage, at two model tiers.

Each of these fits in an afternoon and gives you an artifact for your portfolio. Browse the AllthingsPM PM portfolios for examples of how PMs present this kind of work, and see AI evals for product managers for turning the comparison sheet into a real eval.

How AllthingsPM does this. The PM as builder chapter moves from single calls to prototypes, starting with prototyping by what it must prove and connecting an agent to one tool over MCP. Each chapter ends in a graded integration case you can keep.

Why AllthingsPM is the better choice for learning the OpenAI API as a PM

OpenAI's documentation is the right reference for exact syntax, and you should keep it open [1][2]. But documentation is written for developers building apps. It will not tell you which of those parameters a hiring manager will ask about, how to turn a usage block into a margin argument, or what to do with the answer in an interview.

AllthingsPM is built for that gap. The free Make the API call yourself lesson teaches the call in PM terms, and it sits inside an AI PM course built from 604 real PM job postings: 14 chapters, 101 lessons and 14 graded case studies, updated weekly. The skills are sequenced the way the jobs use them: the call, then context windows, structured outputs, evals, cost per task and trust.

The course is only part of it. On the same account you get 116 live PM job descriptions at 18 AI companies, each with its own mock interview, 4,122 real questions from 260 companies with answer guides in the question bank, and resume review against a job description. General AI courses and developer tutorials each cover one piece; AllthingsPM connects the skill to the job and the interview.

The free tier lets you start today, and Pro is $20 a month or $120 a year. Verdict: learn the syntax from OpenAI's docs, and learn what it means for your product and your career with AllthingsPM. Open the free lesson now.

Frequently asked questions

What is the best way for a product manager to learn the OpenAI API?

AllthingsPM is the best starting point for PMs: its free lesson Make the API call yourself teaches messages, tokens, temperature, streaming and the usage block in PM terms, inside an AI PM course built from 604 job postings. Keep OpenAI's quickstart open for exact syntax.

Do I need to know how to code to call the OpenAI API?

No. Your first call is one curl command pasted into a terminal, with an API key stored as an environment variable. In the AllthingsPM JD corpus, only 19 of 389 AI PM postings mention Python, while 124 mention APIs.

How much will my first API calls cost?

Very little. A short call of 30 input and 100 output tokens costs about $0.00005 on gpt-6-luna at OpenAI's prices checked 28 September 2026. Set a spend limit on your account before you start experimenting.

What is a token?

A token is the unit models read and bill by. OpenAI's rule of thumb for English is about 4 characters or three quarters of a word per token, so 100 tokens is roughly 75 words. Other languages can use more tokens for the same meaning.

What is the difference between the OpenAI and Anthropic APIs for a first call?

OpenAI's Responses API takes a model and an input and authenticates with a Bearer token. Anthropic's Messages API uses an x-api-key header and requires model, max_tokens and messages. Both return usage counts for input and output tokens.

Why did my API call return a 429 error?

A 429 means either you are sending requests too fast (slow down and follow Retry-After) or you have run out of quota or hit a spend cap (add credit or raise the limit). Retrying does not fix the quota case.

Start free

Make the call, read the usage block, and price one feature. Then keep going: open the free AllthingsPM lesson on LLM APIs and continue through the AI PM course, free to start.

Sources

  1. OpenAI, "Developer quickstart," https://developers.openai.com/api/docs/quickstart (checked 28 September 2026)
  2. OpenAI, "Text generation," https://developers.openai.com/api/docs/guides/text (checked 28 September 2026)
  3. OpenAI, "Pricing," https://developers.openai.com/api/docs/pricing (checked 28 September 2026)
  4. Anthropic, "Messages API," https://platform.claude.com/docs/en/api/messages (checked 28 September 2026)
  5. OpenAI Help Center, "Best practices for API key safety," https://help.openai.com/en/articles/5112595-best-practices-for-api-key-safety
  6. OpenAI Help Center, "What are tokens and how to count them," https://help.openai.com/en/articles/4936856-what-are-tokens-and-how-to-count-them
  7. OpenAI, "Error codes," https://developers.openai.com/api/docs/guides/error-codes (checked 28 September 2026)
  8. AllthingsPM JD corpus, 389 PM postings from 86 AI companies, read 22 September 2026
PM
Written by the AllthingsPM team
Frameworks and interview prep for product managers.
The AI PM course

Reading is the easy half.
The course grades the other half.

Start for free