If you use an AI model through an API — to power a chatbot, summarize documents or automate emails — you don’t pay per question. You pay per token. Understanding tokens is the difference between a predictable bill and an unpleasant surprise.
What is a token?
A token is a small piece of text that a language model reads or writes. It can be a short word, part of a longer word, a number or a punctuation mark. Models split all text into tokens before they process it.
As a rough rule for English, one token is about three-quarters of a word, so 1,000 words is roughly 1,300 tokens. The exact split depends on the model’s tokenizer. Other languages, code and unusual words often use more tokens per word.
Input tokens vs. output tokens
Almost every provider prices tokens in two buckets:
- Input tokens: everything you send, including your instructions, the user’s question, and any documents or chat history you include.
- Output tokens: everything the model writes back.
Output tokens usually cost several times more than input tokens, because generating text takes more computing work than reading it. That has a practical consequence: a long answer can cost more than a long question.
Why prices are quoted per million tokens
A single request costs a fraction of a cent, so providers quote prices per million tokens. To get the cost of one request, divide your token count by one million and multiply by the price:
cost per request = (input tokens ÷ 1,000,000 × input price) + (output tokens ÷ 1,000,000 × output price)
A worked example
Imagine a customer support chatbot. Each conversation sends about 1,200 input tokens (instructions, the customer’s message and some context) and gets back about 400 output tokens. Suppose the model charges $3 per million input tokens and $15 per million output tokens. These are illustrative numbers, not any specific provider’s rates.
- Input: 1,200 ÷ 1,000,000 × $3 = $0.0036
- Output: 400 ÷ 1,000,000 × $15 = $0.0060
- Total per request: $0.0096
That looks tiny. But at 500 conversations a day for 30 days, it adds up to about $144 a month. Volume is what turns fractions of a cent into real money.
Hidden costs people forget
- Chat history. In a conversation, each new message usually resends the earlier messages as input, so later turns cost more than early ones.
- System instructions. A long set of instructions is paid for on every single request.
- Retrieved documents. If your app pastes in search results or files, those count as input tokens too.
- Retries. Failed or repeated calls are usually still billed.
- Reasoning tokens. Some models work through a problem before answering and bill that hidden reasoning as output.
Discounts that can lower the bill
Many providers offer cheaper rates for repeated input (often called prompt caching) and for non-urgent jobs processed in batches. Smaller models in the same family are also much cheaper and handle simple tasks well. Check your provider’s pricing page for what’s available, since the options change often. For more ideas, read How to Cut Your AI API Costs.
Estimate your own costs
You don’t need a spreadsheet. Our AI API Cost Calculator lets you enter your token counts, request volume and your model’s current prices to see the cost per request, per day, per month and per year.
Key takeaways
- You pay per token, not per question.
- Output usually costs more than input.
- Chat history, instructions and documents all count as input.
- Multiply by your real volume before deciding a model is cheap.
Leave a Reply