Loading the calculator. If it does not appear, refresh the page or check that JavaScript is enabled.

Estimate model API costs for your AI agent.

One task can trigger many billable model calls. Estimate token costs as context grows, tools return data, and caching changes input prices.

Agent Workload

Choose a usage profile or configure how many sessions and tool calls your agent runs.

Separate tasks or workflows, each with its own context.
Your initial request and follow-up messages in each run.
Searches, API calls, browser actions, code execution, and similar steps. Not sure? Keep the profile default.
Runtime & caching 22 active days · Cache on
Token size assumptions 6 parameters · Profile defaults

Model & Rates

Choose standard model API rates and a market, or enter custom per-million token prices.

USD per 1M tokens
Cache & context assumptions Context, TTL, minimum cache size
Estimated Monthly Cost
/ month
— per agent run

Model API token costs in USD only. Excludes external tools, hosting, token-hour storage and taxes.

Agent runs
Model requests
Peak / run
With caching
Without caching
You save
Token type Cost Share
Fresh inputAwaiting calculation
Cache readAwaiting calculation
Cache writeAwaiting calculation
Output + reasoningAwaiting calculation
See how this simulation works
Request-by-request context growth and billed-token audit for one agent run

Context Growth Across One Agent Run

Context tokens Compaction event
Run simulation to view context growth

Cumulative Billed Tokens Across One Agent Run · Log scale

Fresh input Cache read Cache write Output + reasoning

Cumulative Cost Across One Agent Run

Total Fresh input Cache read Cache write Output + reasoning

Request-Level Execution Trace

Input is the new or rewritten content triggering this request; Context is the full prompt billed by the model. Output includes reasoning.

Turn # Request Input Response Output Context Fresh input Cache read Cache write Cost
How it works

Why Agent Costs Compound Across a Session

An agent can make several model requests for one user task. Each request can resend the growing conversation, add tool results, and receive a new response. The simulator models that request loop instead of multiplying one prompt by a token price.

The AI Agent Request Loop
1. User Prompt User Goal User asks the agent to research, create, analyze, or complete a task.
2. Model Plan Tool Call Emitted Model reasons and outputs parameters for a search, API, browser, or other tool.
3. Execution Tool Output Collected The environment executes the tool and returns data, an action result, or an error.
4. Resend Context Prompt Rebuilt History and tool results return to the model. Eligible prefixes may be reused from cache.
5. Next Step or Reply Iterate or Resolve Loop repeats until goal is satisfied, then model writes final answer.
Read the full methodology →
Modeling Scope

Key Assumptions

These boundaries matter most when interpreting the estimate.

View all assumptions →
Common Questions

AI Agent Cost FAQ

Why does one agent task make multiple model requests?

Planning, parallel tool calls, tool results, evaluation, and the final reply each participate in an iterative request loop.

What is Cache read?

Cache read is a reused prompt prefix billed at the provider's cache-hit rate instead of the ordinary input rate.

Why do some models have no Cache write price?

Some providers populate caches automatically. Their cache misses bill as Fresh input rather than through a separate Cache write meter.

Why does context grow during a run?

The agent resends prior messages and adds tool results or other new content until compaction reduces the retained history.

How accurate is the estimate?

It is a request-level scenario estimate. Actual bills depend on real token sizes, cache behavior, retries, routing, and vendor pricing rules.

Share your estimate

Created locally in your browser. Review the estimate before saving or sharing; no image or calculator inputs are uploaded to generate it.

The QR opens , not your saved settings. Use a copied result link to reproduce the detailed configuration.