Find the best llm apis for your workload
Start with the priced options, then check capabilities, limits and commitments. We do not score qualities we have not measured.
What this comparison covers
Commercial API inference for generative language and reasoning models. Shared multi-tenant API endpoints, standard rate limits, and pay-as-you-go billing without committed use discounts or provisioned throughput. Fine-tuning, batch API discounts, embeddings, and audio generation are excluded.
Evaluating 4 products from 4 published in this category. USD monthly equivalents, before tax. Features and service compatibility need separate evaluation.
- Standard input tokens
- 10,000,000 tokens / month
- Output tokens
- 2,000,000 tokens / month
- Cached prompt tokens
- 5,000,000 tokens / month
Choose the fit, then compare the bill
- Confirm the product supports your requirements. “Not modeled” does not mean a vendor cannot support them.
- Compare included usage, hard limits and paid extras at your expected volume.
- Account for annual commitments, migration effort and the exclusions shown below.
Priced options to evaluate
2 eligible products. Quote-only, expired and unsupported configurations are excluded from cost rankings.
OpenAI API
Industry-standard foundation models with automatic prompt caching and vision.
$3.08 / month
GPT-4o mini
Official pricing · Checked 2026-10-06
Cost breakdown and exclusions
- Monthly subscription: $0
- Standard input tokens: $1.50
- Output completion tokens: $1.20
- Cached prompt tokens: $0.38
- Automatic 50% prompt caching discount applies to prompts with 1,024+ identical leading tokens.
- Fine-tuning, Batch API discounts, audio/realtime APIs, and image generation.
GPT-4o mini: $3.08/month at reference usage.
Groq Cloud
LPU Inference Engine delivering unmatched token output speeds.
$10.43 / month
Llama 3.3 70B (Groq)
Official pricing · Checked 2026-10-06
Cost breakdown and exclusions
- Monthly subscription: $0
- Standard input tokens: $5.90
- Output completion tokens: $1.58
- Cached prompt tokens: $2.95
- Industry-leading 250+ tokens per second inference speeds on dedicated LPU hardware.
- Context prompt caching discounts, vision inputs on this model.
Llama 3.3 70B (Groq): $10.43/month at reference usage.
Questions before you decide
Which LLM APIs product is cheapest?
OpenAI API has a lowest modeled cost among the eligible products shown at reference usage. Ties share a rank. This is not a claim about every vendor or workload.