How to budget AI tokens without mixing up volume and price

6 October 2026 · Hypothetical examples

A price per million tokens does not tell you what a month will cost. You also need requests, days of use, and input and output tokens per request. Keeping these quantities separate reveals what changes when usage grows or responses get longer.

Every amount in this guide is a hypothetical USD scenario. The USD 2 input and USD 8 output rates per million do not represent any model’s official prices.

Open the token cost calculator

A budget with two components

Period cost = requests/day × days × (input tokens × input rate + output tokens × output rate) ÷ 1,000,000.

Assume 100 requests a day for 30 days. Each request uses 1,000 input tokens and generates 500 output tokens. That is 3 million input and 1.5 million output tokens for the period: USD 6 + USD 12 = USD 18. Output uses fewer tokens, yet costs more at these hypothetical rates.

What changes when requests increase

With all other assumptions held constant, cost scales in proportion to requests. The chart compares 100, 300 and 1,000 requests a day. It is not a forecast of users or actual usage: one action in an application can trigger several API calls.

Bar chart: 100, 300 and 1,000 daily requests cost USD 18, 54 and 180 over 30 days, with separate input and output costs at hypothetical rates.
Hypothetical example: 30 days, 1,000 input and 500 output tokens per request; USD 2/8 per million. Totals: USD 18, 54 and 180. Segment length represents cost.
Request-volume example data (USD)
Requests/dayInput USDOutput USDTotal USD
10061218
300183654
100060120180

Output length changes the budget too

At 100 daily requests and 1,000 input tokens, increasing output from 250 to 1,000 tokens moves total cost from USD 12 to USD 30. Doubling output from 500 to 1,000 tokens does not double the total: input still costs USD 6. Review the two components separately.

Bar chart: 250, 500 and 1,000 output tokens per request give hypothetical monthly totals of USD 12, 18 and 30, with input fixed at USD 6.
Hypothetical sensitivity at 100 requests/day for 30 days. Input is fixed at 1,000 tokens and USD 6. Output is 250/500/1,000 tokens; total is USD 12/18/30. This does not assume equal response quality.
Output-sensitivity data (USD)
Output tokens/requestInput USDOutput USDTotal USD
2506612
50061218
100062430

Turn a scenario into a useful budget

Measure requests and tokens over a window that represents your usage, separating workdays, tests and peaks. A conversation may resend history and increase input on each turn. Do not convert words or characters to tokens at a fixed ratio: language, content and tokenizers can change the relationship.

  1. Record the date, model and unit for each provider rate.
  2. Use token counts and calls from your application; document the window you extrapolate.
  3. Calculate a base case and a higher-usage case without inventing a growth rate.
  4. Enter the assumptions in the existing calculator, then reconcile them with actual billing details.

What this calculation leaves out

This simple model excludes taxes, currency conversion, caching, batch, tools, images, audio, retries and fixed fees. Long contexts, reasoning, tiers or minimums may require separate components. Do not add a contingency percentage without explaining which risk it covers. Record additional charges separately and check their units.

Sources and method

Claude Platform’s pricing documentation separates input, output, caching and other modes. It is linked as an example of billing structure, not a recommendation or the source of this guide’s fictional rates. Consulted on 6 October 2026.

Official Claude Platform pricing documentation

Original CosteSoftware charts calculated from the visible assumptions. The tables provide the same data without relying on images. The linked calculator uses editable rates and does not connect to an API.

Open the token cost calculator