LLM APIs: cheapest modeled plans
Eligible plans ordered by monthly equivalent at the workload below. A low subscription fee alone does not establish the cheapest total.
What this comparison covers
Commercial API inference for generative language and reasoning models. Shared multi-tenant API endpoints, standard rate limits, and pay-as-you-go billing without committed use discounts or provisioned throughput. Fine-tuning, batch API discounts, embeddings, and audio generation are excluded.
Evaluating 4 products from 4 published in this category. USD monthly equivalents, before tax. Features and service compatibility need separate evaluation.
- Standard input tokens
- 10,000,000 tokens / month
- Output tokens
- 2,000,000 tokens / month
- Cached prompt tokens
- 5,000,000 tokens / month
Cost ranking at this workload
2 eligible products. Quote-only, expired and unsupported configurations are excluded from cost rankings.
OpenAI API
Industry-standard foundation models with automatic prompt caching and vision.
$3.08 / month
GPT-4o mini
Official pricing · Checked 2026-10-06
Cost breakdown and exclusions
- Monthly subscription: $0
- Standard input tokens: $1.50
- Output completion tokens: $1.20
- Cached prompt tokens: $0.38
- Automatic 50% prompt caching discount applies to prompts with 1,024+ identical leading tokens.
- Fine-tuning, Batch API discounts, audio/realtime APIs, and image generation.
GPT-4o mini: $3.08/month at reference usage.
Groq Cloud
LPU Inference Engine delivering unmatched token output speeds.
$10.43 / month
Llama 3.3 70B (Groq)
Official pricing · Checked 2026-10-06
Cost breakdown and exclusions
- Monthly subscription: $0
- Standard input tokens: $5.90
- Output completion tokens: $1.58
- Cached prompt tokens: $2.95
- Industry-leading 250+ tokens per second inference speeds on dedicated LPU hardware.
- Context prompt caching discounts, vision inputs on this model.
Llama 3.3 70B (Groq): $10.43/month at reference usage.
Questions before you decide
Which LLM APIs product is cheapest?
OpenAI API has a lowest modeled cost among the eligible products shown at reference usage. Ties share a rank. This is not a claim about every vendor or workload.