COSTS YOU CAN INSPECT

LLM APIs: cheapest modeled plans

Eligible plans ordered by monthly equivalent at the workload below. A low subscription fee alone does not establish the cheapest total.

What this comparison covers

Commercial API inference for generative language and reasoning models. Shared multi-tenant API endpoints, standard rate limits, and pay-as-you-go billing without committed use discounts or provisioned throughput. Fine-tuning, batch API discounts, embeddings, and audio generation are excluded.

Evaluating 4 products from 4 published in this category. USD monthly equivalents, before tax. Features and service compatibility need separate evaluation.

Standard input tokens
10,000,000 tokens / month
Output tokens
2,000,000 tokens / month
Cached prompt tokens
5,000,000 tokens / month
Change the workload and compare plans →

Cost ranking at this workload

2 eligible products. Quote-only, expired and unsupported configurations are excluded from cost rankings.

OpenAI API

Industry-standard foundation models with automatic prompt caching and vision.

$3.08 / month

GPT-4o mini

Official pricing · Checked 2026-10-06

Cost breakdown and exclusions
  • Monthly subscription: $0
  • Standard input tokens: $1.50
  • Output completion tokens: $1.20
  • Cached prompt tokens: $0.38
  • Automatic 50% prompt caching discount applies to prompts with 1,024+ identical leading tokens.
  • Fine-tuning, Batch API discounts, audio/realtime APIs, and image generation.

GPT-4o mini: $3.08/month at reference usage.

Groq Cloud

LPU Inference Engine delivering unmatched token output speeds.

$10.43 / month

Llama 3.3 70B (Groq)

Official pricing · Checked 2026-10-06

Cost breakdown and exclusions
  • Monthly subscription: $0
  • Standard input tokens: $5.90
  • Output completion tokens: $1.58
  • Cached prompt tokens: $2.95
  • Industry-leading 250+ tokens per second inference speeds on dedicated LPU hardware.
  • Context prompt caching discounts, vision inputs on this model.

Llama 3.3 70B (Groq): $10.43/month at reference usage.

Questions before you decide

Which LLM APIs product is cheapest?

OpenAI API has a lowest modeled cost among the eligible products shown at reference usage. Ties share a rank. This is not a claim about every vendor or workload.