Skip to content
OperatorNest

Estimate AI API costs by model

Estimate API token spend for several models using published input and output rates.

Example: Priya at Brackenwold Studio: 10,000 input tokens per request, including 5,000 cached tokens; 2,000 output tokens; 100 requests a day and 3,000 a month.

Example: 10,000 input tokens per request, 5,000 cached, and 2,000 output tokens.

Anthropic
OpenAI
xAI
Alibaba
DeepSeek
Google
Meta
Mistral
Moonshot AI
Cohere
Amazon Bedrock
Estimated request, daily, and monthly costs by selected model
ModelPer requestPer dayPer monthBatch
Claude Haiku 4.5 (latest)Official sourceChecked September 28, 2026$0.0155$1.55$46.5Standard
Claude Haiku 4.5Official sourceChecked September 28, 2026$0.0155$1.55$46.5Standard
Gemini 3.5 Flash LiteOfficial sourceChecked September 28, 2026$0.0067$0.665$19.95Standard
Gemini Flash-Lite LatestOfficial sourceChecked September 28, 2026$0.0067$0.665$19.95Standard
Gemini 3.1 Flash LiteOfficial sourceChecked September 28, 2026$0.0044$0.4375$13.125Standard
Gemini 3.1 Flash Lite PreviewOfficial sourceChecked September 28, 2026$0.0044$0.4375$13.125Standard

Estimates use the catalog snapshot shown above. Read each linked provider page for rate terms. Rates exclude non-token charges. Model data from models.dev (MIT).

Compare model rates and reviewed changes.

Embed this tool

Copy this snippet to show the tool on your site.

How it works

Each request uses input tokens × input price plus output tokens × output price, divided by one million. Published prompt-length price tiers are applied based on total input tokens.

Cached input tokens replace the same number of regular input tokens when the provider publishes a cache rate. Batch discounts apply only to models with a published discount; provider cache pricing can differ in batch mode.

Monthly requests are entered separately, so you can estimate an uneven month without assuming 30 days.

Sources and checked dates

Limits

  • This estimates text-token charges only. It leaves out tools, images, audio, taxes, credits, regional adjustments, and provider-specific charges.
  • Published long-context rates are applied for OpenAI, Gemini 3.1 Pro, and Grok 4.7. DeepSeek peak rates are used; prompt size and time of day can change provider charges.
  • Qwen rates are published in CNY and host prices vary for open-weight Llama models; neither is included in USD totals.

Common questions

How is the estimate calculated?
Input and output tokens are multiplied by each selected model’s published rate per million tokens. Cached tokens use the published cache rate when one is listed. Published prompt-length tiers and batch rates are applied when available.
What does the batch option change?
It applies a provider-published batch discount to eligible models. Batch work runs asynchronously, so the result may not arrive with an interactive request.
Why can’t I estimate every model?
Some prices vary by region, input length, or host. This page leaves out non-USD rates and unverified prices instead of converting or guessing.

OperatorNest can take on repeat work, check with you before consequential steps, and leave a receipt for the result. See how an always-on operator works.

Hand off your first task tonight.

Tell us your email and what you'd hand off first. We'll send your access details and help you set up your operator.