Estimate AI API costs by model
Estimate API token spend for several models using published input and output rates.
Example: 10,000 input tokens per request, 5,000 cached, and 2,000 output tokens.
| Model | Per request | Per day | Per month | Batch |
|---|---|---|---|---|
| Claude Haiku 4.5 (latest)Official source | $0.0155 | $1.55 | $46.5 | Standard |
| Claude Haiku 4.5Official source | $0.0155 | $1.55 | $46.5 | Standard |
| Gemini 3.5 Flash LiteOfficial source | $0.0067 | $0.665 | $19.95 | Standard |
| Gemini Flash-Lite LatestOfficial source | $0.0067 | $0.665 | $19.95 | Standard |
| Gemini 3.1 Flash LiteOfficial source | $0.0044 | $0.4375 | $13.125 | Standard |
| Gemini 3.1 Flash Lite PreviewOfficial source | $0.0044 | $0.4375 | $13.125 | Standard |
Estimates use the catalog snapshot shown above. Read each linked provider page for rate terms. Rates exclude non-token charges. Model data from models.dev (MIT).
Embed this tool
Copy this snippet to show the tool on your site.
How it works
Each request uses input tokens × input price plus output tokens × output price, divided by one million. Published prompt-length price tiers are applied based on total input tokens.
Cached input tokens replace the same number of regular input tokens when the provider publishes a cache rate. Batch discounts apply only to models with a published discount; provider cache pricing can differ in batch mode.
Monthly requests are entered separately, so you can estimate an uneven month without assuming 30 days.
Sources and checked dates
Limits
- This estimates text-token charges only. It leaves out tools, images, audio, taxes, credits, regional adjustments, and provider-specific charges.
- Published long-context rates are applied for OpenAI, Gemini 3.1 Pro, and Grok 4.7. DeepSeek peak rates are used; prompt size and time of day can change provider charges.
- Qwen rates are published in CNY and host prices vary for open-weight Llama models; neither is included in USD totals.
Common questions
- How is the estimate calculated?
- Input and output tokens are multiplied by each selected model’s published rate per million tokens. Cached tokens use the published cache rate when one is listed. Published prompt-length tiers and batch rates are applied when available.
- What does the batch option change?
- It applies a provider-published batch discount to eligible models. Batch work runs asynchronously, so the result may not arrive with an interactive request.
- Why can’t I estimate every model?
- Some prices vary by region, input length, or host. This page leaves out non-USD rates and unverified prices instead of converting or guessing.
Related tools
- AI token counter
Count tokens in pasted text and estimate the input cost for a selected model.
- Context window calculator
Estimate model context use before sending a document and planned output.
- AI subscription cost calculator
Compare the cost and listed capabilities of AI subscriptions.
OperatorNest can take on repeat work, check with you before consequential steps, and leave a receipt for the result. See how an always-on operator works.
Hand off your first task tonight.
Tell us your email and what you'd hand off first. We'll send your access details and help you set up your operator.