Estimate your AI feature's model bill before you build it
The Monthly Model Bill
Most AI MVP budgets I see cover the build and forget the monthly line. That line is the model bill, and at a few thousand users it can pass what the build cost. One formula gets you a usable estimate.
The Formula
monthly cost = requests per month × (input tokens × input price + output tokens × output price)
Requests per month is users times how often they use the feature.
Input tokens include everything you send: system prompt, retrieved context, chat history. Resending all of that on every request is why input usually dominates.
Worked Example
- 10,000 monthly users
- 40 requests each
- 4,000 input and 400 output tokens per request
That's 400,000 requests, 1.6B input tokens and 160M output tokens a month.
Model Cost Comparison (list price per 1M tokens)
| Model | Input | Output | Monthly |
|---|---|---|---|
| Claude Sonnet 5 ($2 / $10) | $3,200 | $1,600 | $4,800 |
| Gemini 3.1 Flash-Lite ($0.25 / $1.50) | $400 | $240 | $640 |
Same feature, 7.5x apart.
Routing and Caching
Routing
Most requests don't need the big model. Send the easy 80% to Flash-Lite and the hard 20% to Sonnet, and the $4,800 bill drops to about $1,470.
Caching
If 3,000 of those 4,000 input tokens are the same system prompt and reference docs every time, cache them. Gemini 3.1 Flash-Lite bills cached input at $0.025 per 1M instead of $0.25, which takes the Flash-Lite bill from $640 to about $370 a month. Storage is $0.50 per 1M tokens per hour, so keeping a 3,000-token cache alive all month adds about $1.
Both are build decisions. Adding them after launch means rewriting the request path.
What I'd Do Before Writing Code
- Write down users, requests per user and token counts per request, even rough ones.
- Run the formula for a big and a small model.
- If the big-model number scares you, design routing and caching in from day one.
If you want the build cost and the run cost side by side, our AI product cost estimator does both, including the monthly bill at your user count. I also wrote up how we scope and price AI product development in checkpoints.
Prices are list prices from the Anthropic and Google pricing pages, October 2026.
Comments
No comments yet. Start the discussion.