Newtuple
Back to Blog
LLM Engineering

LLM Pricing Cheat Sheet: Compare Costs & Usage Tiers

Quickly compare pricing models and tiers for top large-language models. Find the best cost-effective plan for your project

August 18, 202614 min readUpdated August 20, 2026
LLM Pricing Cheat Sheet: Compare Costs & Usage Tiers

Last update - 20 August 2026

Latest updates: Current flagship, value, embedding, and fine-tuning prices from official provider pages

How much does it cost for you to integrate your application with a large language model? Here's an article that will show the rough relative costs of using different models for chat / completion models, costs for embedding and for making use of a fine-tuned model. I will keep this page updated, so bookmark this and use this as a reference guide whenever you're working on your next cool idea. Happy building!

Additionally, if you have information on other models OR if you have additional pricing inputs to share, do write in the comments.

A quick intro before moving ahead with the pricing comparison

Interest and usage of large language models have exploded. Most teams have already tried general-purpose AI assistants, and the next value unlock is to apply a model to a specific use case and application. The 2 key questions to answer for you are:

1. Which model should I choose?

2. How much will it cost me / my organization?

While I cannot answer question 1 in this article (the short answer is: it depends), we can certainly look into question 2 to understand how much it costs, in order for us to make a better business case for the development of a custom app. If you're looking for some help with assessing which model to use, its solution design or development of your LLM based use case, fill up this form and let's talk!

Tokens vs Characters vs Everything else

The main challenge in comparing providers is that the headline token rate is no longer the full price. Providers can charge different rates for input, cached input, output, long context, tools, search, regions, data residency, and service tiers. The tables below normalize text prices to US dollars per million tokens where possible, and call out important extra charges.

Some definitions:

1. Session - I created the concept of a "session" to make it easier for us to visualize costs for an application we're building out. The definition of 1 session is: A session between a human and your bot consisting of 100 words x 5 characters each = 500 characters = 125 tokens (in english language). Assume 25% of the tokens are consumed by a human and 75% is consumed by your bot generating outputs.

2. Page - I like to think about embeddings in terms of pages rather than tokens. In this sheet, I've assumed 1 page = 500 words x 5 characters = 625 tokens per page

Pricing

So let's get into it!

A. Usage of General Prompt / Completion or Chat models

Let's first look at costs for all completion and chat models, the ones that we would use for most often: "ChatGPT for my App", chatbots, knowledge retrieval bots (+ add costs of embeddings to this)

ProviderModelContextPrice / 1M input tokensCached inputPrice / 1M output tokensNotes
OpenAIGPT-5.6 Sol1.05M$5.00$0.50$30.00Standard API rate
OpenAIGPT-5.6 Terra1.05M$2.00$0.20$12.00Standard API rate
OpenAIGPT-5.6 Luna1.05M$0.20$0.02$1.20Standard API rate
OpenAIGPT-5.51.05M$5.00$0.50$30.00Requests above 272k input tokens use higher long-context rates
OpenAIGPT-5.5 Pro1.05M$30.00$180.00Pro model
OpenAIGPT-5.41.05M$2.50$0.25$15.00Requests above 272k input tokens use higher long-context rates
OpenAIGPT-5.4 mini400k$0.75$0.075$4.50Standard API rate
OpenAIGPT-5.4 nano400k$0.20$0.02$1.25Standard API rate
OpenAIGPT-5.3 Codex400k$1.75$0.175$14.00Specialized coding model
OpenAIGPT-5.2400k$1.75$0.175$14.00Previous model
OpenAIGPT-5.2 Pro400k$21.00$168.00Previous Pro model
OpenAIGPT-5400k$1.25$0.125$10.00Previous model
OpenAIGPT-5 mini400k$0.25$0.025$2.00Previous mini model
OpenAIGPT-5 nano400k$0.05$0.005$0.40Previous nano model
OpenAIGPT-5 Pro400k$15.00$120.00Previous Pro model
OpenAIo3200k$2.00$0.50$8.00Reasoning model
OpenAIo3-pro200k$20.00$80.00Pro reasoning model
OpenAIGPT-4.11.05M$2.00$0.50$8.00Previous non-reasoning model
OpenAIGPT-4.1 mini1.05M$0.40$0.10$1.60Previous non-reasoning model
OpenAIGPT-4.1 nano1.05M$0.10$0.025$0.40Previous non-reasoning model
OpenAIGPT-4o128k$2.50$1.25$10.00Previous multimodal model
OpenAIGPT-4o mini128k$0.15$0.075$0.60Previous multimodal model
AnthropicClaude Fable 51M$10.00$1.00$50.00Standard Claude API rate
AnthropicClaude Mythos 51M$10.00$1.00$50.00Limited availability
AnthropicClaude Opus 51M$5.00$0.50$25.00Standard mode; fast mode is $10 / $50
AnthropicClaude Opus 4.81M$5.00$0.50$25.00Standard Claude API rate
AnthropicClaude Opus 4.71M$5.00$0.50$25.00Standard Claude API rate
AnthropicClaude Opus 4.6See model card$5.00$0.50$25.00Standard Claude API rate
AnthropicClaude Opus 4.5See model card$5.00$0.50$25.00Standard Claude API rate
AnthropicClaude Opus 4.1See model card$15.00$1.50$75.00Deprecated
AnthropicClaude Opus 4See model card$15.00$1.50$75.00Retired except on Google Cloud
AnthropicClaude Sonnet 51M$2.00$0.20$10.00Promotional rate through 31 August 2026; then $3 / $15
AnthropicClaude Sonnet 4.6See model card$3.00$0.30$15.00Standard Claude API rate
AnthropicClaude Sonnet 4.5See model card$3.00$0.30$15.00Standard Claude API rate
AnthropicClaude Sonnet 4See model card$3.00$0.30$15.00Retired except on Bedrock and Google Cloud
AnthropicClaude Haiku 4.5200k$1.00$0.10$5.00Standard Claude API rate
AnthropicClaude Haiku 3.5See model card$0.80$0.08$4.00Retired except on Bedrock and Google Cloud
Google Vertex AIGemini 3.1 Pro PreviewTiered$2.00$0.20$12.00Global rate for requests up to 200k input tokens; long-context input/output is $4 / $18
Google Vertex AIGemini 3.7 FlashTiered$0.75$0.075$3.75Introductory global rate through 31 December 2026
Google Vertex AIGemini 3.6 FlashTiered$0.75$0.075$3.75Introductory global rate through 31 December 2026
Google Vertex AIGemini 3.5 FlashTiered$1.50$0.15$9.00Global standard rate
Google Vertex AIGemini 3.5 Flash-LiteTiered$0.30$0.03$2.50Global standard rate
Google Vertex AIGemini 3 Flash PreviewTiered$0.50$0.05$3.00Text, image, and video input rate shown
Google Vertex AIGemini 3.1 Flash-LiteTiered$0.25$0.025$1.50Text, image, and video input rate shown
Google Vertex AIGemini Omni FlashTiered$1.50$0.15$9.00Multimodal global rate
Google Vertex AIGemma 4 26BSee model card$0.15$0.015$0.60Open-model endpoint on Vertex AI
xAIGrok 4.5500k$2.00$0.30$6.00Long-context rates from 200k tokens are $4 / $0.60 / $12
xAIGrok 4.31M$1.25$0.20$2.50Long-context rates from 200k tokens are $2.50 / $0.40 / $5
xAIGrok 4.20 Reasoning1M$1.25$0.20$2.50Model ID grok-4.20-0309-reasoning; long-context rates are 2×
xAIGrok 4.20 Non-Reasoning1M$1.25$0.20$2.50Model ID grok-4.20-0309-non-reasoning; long-context rates are 2×
xAIGrok 4.20 Multi-Agent1M$1.25$0.20$2.50Beta multi-agent model; long-context rates are 2×
xAIGrok Build 0.1256k$1.00$0.20$2.00Early-access coding model; long-context rates from 200k tokens are 2×
MistralMistral Medium 3.5See model card$1.50$0.15$7.50Direct Mistral API; cached input uses the stated 90% discount
MistralMistral Small 4See model card$0.15$0.015$0.60Direct Mistral API
MistralMistral Large 3See model card$0.50$0.05$1.50Direct Mistral API
MistralMinistral 3 14BSee model card$0.20$0.02$0.20Direct Mistral API
MistralMinistral 3 8BSee model card$0.15$0.015$0.15Direct Mistral API
MistralMinistral 3 3BSee model card$0.10$0.01$0.10Direct Mistral API
MistralDevstral 2See model card$0.40$2.00Coding model
MistralDevstral Small 2See model card$0.10$0.30Labs coding model
MistralCodestralSee model card$0.30$0.03$0.90Direct Mistral API
MistralMagistral MediumSee model card$2.00$5.00Reasoning model
MistralMagistral SmallSee model card$0.50$1.50Reasoning model
MistralVoxtral SmallSee model card$0.10$0.40Text-token rate; audio input is $0.004 per minute
MistralLeanstral 1.5See model cardFreeFreeFreeExperimental model; limited-period free endpoint
CohereCommand A256k$2.50$10.00Current Command A model family
CohereCommand R+128k$2.50$10.00Model ID command-r-plus-08-2024
CohereCommand R128k$0.15$0.60Model ID command-r-08-2024
CohereAya Expanse 32B128k$0.50$1.50Research model API rate
CohereAya Expanse 8B128k$0.50$1.50Research model API rate
CohereCommand R+ 04-2024128k$3.00$15.00Legacy model
CohereCommand R 03-2024128k$0.50$1.50Legacy model
CohereCommandSee model card$1.00$2.00Legacy model for existing customers
CohereCommand-lightSee model card$0.30$0.60Legacy model for existing customers
PerplexitySonar128k$1.00$1.00A separate $5–$12 per 1,000 requests fee applies
PerplexitySonar Pro200k$3.00$15.00A separate $6–$14 per 1,000 requests fee applies
PerplexitySonar Reasoning Pro128k$2.00$8.00A separate $6–$14 per 1,000 requests fee applies
PerplexitySonar Deep Research128k$2.00$8.00Citation, reasoning-token, and search-query fees also apply
GroqLlama 3.1 8B Instant131k$0.05$0.08Production model
GroqLlama 3.3 70B Versatile131k$0.59$0.79Production model
GroqGPT OSS 120B131k$0.15$0.60Production model
GroqGPT OSS 20B131k$0.075$0.30Production model
GroqQwen 3.6 27B131k$0.60$3.00Preview model
GroqSafety GPT OSS 20B131k$0.075$0.30Preview safeguard model
GroqLlama Prompt Guard 2 22M512$0.03$0.03Preview safeguard model
GroqLlama Prompt Guard 2 86M512$0.04$0.04Preview safeguard model
DeepSeekDeepSeek V4 Flash1M$0.14$0.0028$0.28Cache-miss input rate shown
DeepSeekDeepSeek V4 Pro1M$0.435$0.003625$0.87Cache-miss input rate shown

Prices are in USD and were verified on 20 August 2026. The table now contains 88 priced model rows. Rates are standard direct-API or named global rates unless a note says otherwise. Batch, flex, priority, data-residency, regional, long-context, tool-call, search, and image/audio token charges can change the total. AWS Bedrock and Azure prices can also differ by region and endpoint, so check their calculators before you commit to a budget.

Perplexity currently publishes four first-party Sonar models. DeepSeek currently publishes two direct API models. Cohere also lists Command A+ and Command R7B, but it does not publish a comparable self-serve token rate for them, so they are not shown as priced rows.

Official sources: OpenAI pricing and model catalog, Anthropic pricing, Google Vertex AI pricing, xAI pricing, Mistral API pricing, Cohere pricing, Perplexity pricing, Groq models, DeepSeek pricing, and AWS Bedrock pricing.

  1. For a mixed chat workload, calculate input and output separately. The earlier 25% input / 75% output session split remains a useful rough estimate.

  2. A session of 125 tokens is only a planning unit. System prompts, retrieved context, reasoning tokens, cached prompts, and tool results can make real requests much larger.

B. Embedding Costs

What does it cost for you to embed your knowledge base using one of the embedding models available? Let's look at 3 scenarios - a knowledge base of 10,000 pages, 100,000 pages and 1 million pages.

ProviderModelPrice per 1M input tokens10,000 pages100,000 pages1,000,000 pages
OpenAItext-embedding-3-small$0.02$0.13$1.25$12.50
OpenAItext-embedding-3-large$0.13$0.81$8.13$81.25
Google Vertex AIGemini Embedding$0.15$0.94$9.38$93.75
Google Vertex AIGemini Embedding 2, text input$0.20$1.25$12.50$125.00
MistralMistral Embed$0.10$0.63$6.25$62.50
MistralCodestral Embed$0.15$0.94$9.38$93.75
Perplexitypplx-embed-v1-0.6b$0.004$0.03$0.25$2.50
Perplexitypplx-embed-v1-4b$0.03$0.19$1.88$18.75
Perplexitypplx-embed-context-v1-0.6b$0.008$0.05$0.50$5.00
Perplexitypplx-embed-context-v1-4b$0.05$0.31$3.13$31.25
  1. This estimate uses 625 tokens per page, as defined above. It covers embedding API use only.

  2. Vector database, storage, retrieval, reranking, network, and re-embedding costs are not included.

  3. Google lists separate prices for image, video, and audio inputs to its multimodal embedding model.

C. Cost of Fine-tuning a model + usage:

Fine-tuning prices are now harder to compare directly. Some providers use training-token prices, some use training hours, and others provide custom enterprise quotes. OpenAI is also winding down access to its fine-tuning platform. Use the table below as a planning reference, not as a complete project quote.

ProviderModel or serviceTraining priceInference priceCurrent status / note
OpenAIo4-mini-2025-04-16 reinforcement fine-tuning$100 per training hour$4 input / $16 output per 1M tokensOpenAI is winding down its fine-tuning platform; it is closed to new users
Google Vertex AIGemini 3.5 Flash supervised or reinforcement tuning$10 per 1M training tokens$2.25 input / $13.50 output per 1M tokensTuned inference is 1.5× the $1.50 / $9 base price
Google Vertex AIGemini 3.1 Flash-Lite supervised tuning$3 per 1M training tokens$0.375 input / $2.25 output per 1M tokensTuned inference is 1.5× the $0.25 / $1.50 base price
MistralClassifier API model 3B$1 per 1M training tokens; $4 minimum$0.10 input / $0.10 output per 1M tokens$2 monthly storage per model also applies
AnthropicCustom modelsContact salesContact salesNo public self-serve fine-tuning rate
CohereCustom modelsContact salesContact salesEnterprise customization uses custom pricing

Training cost is not the full cost of a tuned model. Add dataset preparation, evaluation, repeated experiments, model storage, and production inference. Check the provider page again before you start a training job.

Pricing scenario: a current monthly chat estimate

For a simple comparison, assume 100,000 sessions per month and 125 tokens per session. With the earlier 25% input / 75% output split, that is 3.125 million input tokens and 9.375 million output tokens per month.

ModelApproximate monthly model cost
OpenAI GPT-5.6 Luna$11.88
Google Gemini 3.7 Flash, introductory rate$37.50
Anthropic Claude Sonnet 5, August 2026 promotional rate$100.00

These figures do not include system prompts, retrieved documents, reasoning tokens, tool calls, search fees, caching, or application infrastructure. In a retrieval-augmented generation system, embedding 100,000 pages once would add about $1.25 with OpenAI text-embedding-3-small or $9.38 with Google Gemini Embedding, before vector database costs.

Fine-tuning and retrieval are not interchangeable. Retrieval is usually the better fit when facts change often or citations matter. Fine-tuning is usually a better fit when you need repeatable behavior, style, or task-specific patterns. Many production systems use both.

Did you enjoy this article? Have I missed something in the pricing calculations? Drop a note below and get the discussion started!

Stay in the loop

Get new posts, product updates, and research notes once a week.

By subscribing you agree to receive updates from Newtuple. You can unsubscribe anytime.

Ready to build production AI?

Talk to our team about AI agents, data platforms, and GenAI accelerators.

Get in Touch