No. 39 · OCT 2026 · 5 Min Read
An Executive's Guide to AI Pricing
Abstract
AI services bill by the token, the seat, and the run. How each meter works, why output costs five times input, and how to budget per accepted result.
Two dollars buys a million tokens of input on Claude Sonnet 5.5. Ten dollars buys a million tokens of output. Those are the list prices on Anthropic’s pricing page today.1 Most AI budget conversations start and end with numbers like these, and they explain almost nothing about what the line item will be next quarter.
A token price is a unit price. It tells you what one pound of flour costs. It does not tell you how many loaves the bakery sells, how much dough gets thrown out, or whether the new oven uses twice as much flour per loaf. Executives who budget from the rate card get surprised by the volume, and the volume is decided by architecture, not procurement.
The Three Meters
AI services charge through three meters, and most organizations are paying on all three at once.
Seats. A flat monthly fee per person for a chat product or a coding assistant. Predictable, easy to approve, and the cheapest way to find out who on the team actually uses the thing. Seat pricing hides the token bill inside the vendor’s margin. Heavy users are subsidized by light ones.
Tokens. The API meter. Every word the model reads is input. Every word it writes is output. A token is roughly three quarters of an English word. You pay per million of each, and the bill arrives in arrears.
Runs and tools. The newer meter. Agent platforms charge for runtime on top of tokens. Anthropic’s Managed Agents bills $0.08 per session-hour of running time. Web search costs $10 per thousand searches, plus the tokens for every result the model reads.1 Each is a rounding error on one request and a real line item across a fleet.
Seats are a staffing decision. Tokens and runs are a systems decision. They need different owners.
Why Output Costs Five Times More
Across the current Claude lineup, output tokens cost five times input tokens. Same ratio from Haiku to Opus.1 Reading is cheap for a model. Writing is expensive, because each output token is generated one at a time and each one has to consider everything before it.
That ratio should change how you read a proposal. A system that reads a 50-page contract and returns a two-line verdict is mostly input, and cheap. A system that drafts a 50-page report from a two-line prompt is mostly output, and five times pricier per token. Two projects with the same “volume of text” can differ in cost by a large multiple depending on which direction the text flows.
Reasoning makes this worse. Models that think before they answer generate that thinking as output, and you pay for it whether or not anyone reads it.
The Discounts Are Architecture
Vendors publish real discounts. You only get them if the system is built to qualify.
Batch. Work that can wait gets half off. Anthropic’s Batch API discounts input and output by 50%.1 Overnight classification, backfills, report generation, and document review rarely need an answer in two seconds. If they run through the interactive endpoint anyway, someone is paying double for latency nobody asked for.
Caching. When many requests share the same opening material (a long system prompt, a policy manual, a product catalog), the provider can cache it. A cache hit costs a tenth of the normal input price on most Claude models.1 That is a 90% discount on the part of the request that never changes. It requires the shared material to sit in the same place at the start of every request, which is a design choice someone has to make on purpose.
Model choice. The spread between the smallest and largest model in a single vendor’s lineup is roughly tenfold. Sending everything to the biggest model is the default because it is the safest default. Model routing is an evaluation problem: you can only send work to the cheaper model when you can prove it does the job.
None of these show up in a pilot. A pilot runs a few hundred interactive requests on the flagship model with no caching. Then it gets approved, scaled, and nobody revisits the shape.
What Moves the Bill
The rate card is rarely the variable that matters. These are.
| Driver | What it means | Who controls it |
|---|---|---|
| Context per request | How much the model reads every call | Engineering |
| Calls per task | How many model turns one piece of work takes | Engineering |
| Output length | How much the model writes, including reasoning | Engineering and product |
| Retries and failures | Work paid for and thrown away | Engineering and QA |
| Usage growth | How many people or jobs hit the system | The business |
Engineering controls the first four. An agent that takes twelve turns to finish a task, re-reading its whole history each turn, pays for that history twelve times. Inference cost is architecture for exactly this reason. The same feature can cost ten cents or ten dollars per run depending on how much control flow sits in ordinary code versus in the model.
Vendors also change the unit underneath you. Anthropic notes that Claude 4.7 and later models use a tokenizer that produces about 30% more tokens for the same text.1 A model upgrade at the same per-token price can still raise the bill. Read the release notes before signing off on “same price, better model.”
Budget Per Accepted Result
The unit an executive should track is cost per accepted result. What did it cost to get one invoice correctly extracted, one ticket correctly resolved, one contract correctly flagged?
That number folds in everything above: the model, the context, the retries, the human review of the output that was wrong. It also exposes the cheap system that is quietly expensive. A model that costs a third as much and fails twice as often has not saved money once someone has to clean up after it.
To get there, ask four questions of any AI line item:
- What is the unit of work, and how many will we do per month?
- What does one unit cost today, including failures and review?
- Which of these requests could wait for batch, and which share cacheable context?
- Who owns the number, and who gets told when it moves?
The fourth question is the one that gets skipped. Token spend lands on an engineering team’s cloud bill while the value lands on an operations team’s scorecard. Nobody owns the ratio, and using AI is management: someone has to.
Pricing pages change every few months. The meters, the output premium, and the discounts for well-shaped work have stayed stable. Budget around those.