Multi-Model · Cost Per Request · Monthly Estimate
| Model | Input | Output | Context |
|---|---|---|---|
| GPT-4o | $2.50 | $10.00 | 128K |
| GPT-4o-mini | $0.15 | $0.60 | 128K |
| GPT-4 Turbo | $10.00 | $30.00 | 128K |
| Claude 3.5 Sonnet | $3.00 | $15.00 | 200K |
| Claude 3 Haiku | $0.25 | $1.25 | 200K |
| Gemini 1.5 Pro | $1.25 | $5.00 | 2M |
| Gemini 1.5 Flash | $0.075 | $0.30 | 1M |
| Llama 3.1 70B* | $0.35 | $0.40 | 128K |
*Via providers like Together, Fireworks, Groq. Prices as of early 2025 — check providers for current rates.
| Language | Tokens per Word | Tokens per Page |
|---|---|---|
| English | ~1.3 | ~800 |
| Spanish | ~1.5 | ~900 |
| Code | ~2.5 | ~1,500 |
| Chinese/Japanese | ~2.0 | ~1,200 |
1 token ≈ 4 characters or ¾ of a word in English.
| Strategy | Savings |
|---|---|
| Use smaller models for simple tasks | 80-95% cheaper |
| Prompt caching (same prefix) | 50% on input |
| Batch API (non-real-time) | 50% discount |
| Limit output with max_tokens | 20-50% on output |
| Fine-tune for repetitive tasks | Fewer tokens needed |