CHOOSE WITH EVIDENCE

Find the best llm apis for your workload

Start with the priced options, then check capabilities, limits and commitments. We do not score qualities we have not measured.

What this comparison covers

Commercial API inference for generative language and reasoning models. Shared multi-tenant API endpoints, standard rate limits, and pay-as-you-go billing without committed use discounts or provisioned throughput. Fine-tuning, batch API discounts, embeddings, and audio generation are excluded.

Evaluating 4 products from 4 published in this category. USD monthly equivalents, before tax. Features and service compatibility need separate evaluation.

Standard input tokens
10,000,000 tokens / month
Output tokens
2,000,000 tokens / month
Cached prompt tokens
5,000,000 tokens / month
Change the workload and compare plans →

Choose the fit, then compare the bill

  1. Confirm the product supports your requirements. “Not modeled” does not mean a vendor cannot support them.
  2. Compare included usage, hard limits and paid extras at your expected volume.
  3. Account for annual commitments, migration effort and the exclusions shown below.

Priced options to evaluate

2 eligible products. Quote-only, expired and unsupported configurations are excluded from cost rankings.

OpenAI API

Industry-standard foundation models with automatic prompt caching and vision.

$3.08 / month

GPT-4o mini

Official pricing · Checked 2026-10-06

Cost breakdown and exclusions
  • Monthly subscription: $0
  • Standard input tokens: $1.50
  • Output completion tokens: $1.20
  • Cached prompt tokens: $0.38
  • Automatic 50% prompt caching discount applies to prompts with 1,024+ identical leading tokens.
  • Fine-tuning, Batch API discounts, audio/realtime APIs, and image generation.

GPT-4o mini: $3.08/month at reference usage.

Groq Cloud

LPU Inference Engine delivering unmatched token output speeds.

$10.43 / month

Llama 3.3 70B (Groq)

Official pricing · Checked 2026-10-06

Cost breakdown and exclusions
  • Monthly subscription: $0
  • Standard input tokens: $5.90
  • Output completion tokens: $1.58
  • Cached prompt tokens: $2.95
  • Industry-leading 250+ tokens per second inference speeds on dedicated LPU hardware.
  • Context prompt caching discounts, vision inputs on this model.

Llama 3.3 70B (Groq): $10.43/month at reference usage.

Questions before you decide

Which LLM APIs product is cheapest?

OpenAI API has a lowest modeled cost among the eligible products shown at reference usage. Ties share a rank. This is not a claim about every vendor or workload.