LLM Pricing Tracker: API and Subscription Costs
A tracker for leading LLM API token prices and consumer subscriptions, with official links, repo snapshot refreshes, price-history charts, and live benchmark/provider snapshots.
Note
Latest repo snapshot in this build: July 13, 2026. Click Check for newer snapshot below to query GitHub for a fresher snapshot. Each browser is limited to one remote check per day.
This page tracks public pricing from official provider pages for major frontier-model vendors I regularly compare: OpenAI (including GPT-5.6 Sol, Terra, and Luna), Google (Gemini and Gemma), Anthropic, xAI, DeepSeek, Qwen, Moonshot/Kimi, Xiaomi/MiMo, MiniMax, Together AI (including GLM-5.2 and public Llama endpoints), Meta/Llama references, and GitHub Copilot.
A few quick cautions before using the numbers:
- API pricing and consumer subscription pricing are different products.
- Some vendors publish tiered pricing by context length, region, or prompt type.
- When a vendor does not publicly expose a comparable token-billing number, I mark that clearly instead of guessing.
- The charts below use snapshots stored in this repo, including daily-refreshable Artificial Analysis benchmark/provider snapshots and manually curated pricing rows from official vendor pages.
API Token Fees
Representative API pricing for major frontier-model vendors. Numbers are taken from official vendor pricing pages; mixed currencies are kept in the vendor’s quoted currency.
| Vendor | Model / Product | Input | Cached input | Output | Notes | Official |
|---|---|---|---|---|---|---|
| OpenAI | GPT-5.6 Sol (<272K context) | $5.00 | $0.50 | $30.00 | OpenAI's highest-capability GPT-5.6 model. Requests above 272K input tokens are billed at 2x input and 1.5x output rates; regional processing adds 10%. | OpenAI API pricing |
| OpenAI | GPT-5.6 Terra (<272K context) | $2.50 | $0.25 | $15.00 | OpenAI's balanced GPT-5.6 model. Requests above 272K input tokens are billed at 2x input and 1.5x output rates; regional processing adds 10%. | OpenAI API pricing |
| OpenAI | GPT-5.6 Luna (<272K context) | $1.00 | $0.10 | $6.00 | OpenAI's fastest and lowest-cost GPT-5.6 model. Requests above 272K input tokens are billed at 2x input and 1.5x output rates; regional processing adds 10%. | OpenAI API pricing |
| Gemini 3.5 Flash | $1.50 | $0.15 | $9.00 | Current Gemini API paid-tier standard pricing for Gemini 3.5 Flash. Google describes it as its most intelligent model built for speed; context-cache storage is listed separately at $1.00 / 1M tokens per hour. | Gemini API pricing | |
| Gemma 4 | Free of charge | Free of charge | Free of charge | Google's current Gemini Developer API pricing page lists Gemma 4 with free input, output, and context caching, while the paid tier remains unavailable. | Google Gemini Developer API pricing | |
| Anthropic | Claude Fable 5 | $10.00 | $1.00 (cache read) | $50.00 | Anthropic restored global access on July 1, 2026. Five-minute prompt-cache writes are $12.50 / 1M tokens; US-only inference is 1.1x standard pricing. | Claude API pricing |
| Anthropic | Claude Mythos 5 (limited access) | $10.00 | $1.00 (cache read) | $50.00 | Invitation-only Project Glasswing model with the same underlying specifications and pricing as Claude Fable 5. It is not a self-serve API model and requires 30-day data retention. | Anthropic Claude Fable 5 and Mythos 5 announcement |
| Anthropic | Claude Sonnet 5 | $2.00 (introductory) | $0.20 (cache read) | $10.00 (introductory) | Introductory pricing through August 31, 2026; standard pricing becomes $3 input, $0.30 cache read, and $15 output per 1M tokens. Prompt-cache writes are $2.50 during the introductory period. | Claude API pricing |
| DeepSeek | deepseek-v4-flash | $0.14 (cache miss) | $0.0028 (cache hit) | $0.28 | Official DeepSeek V4 Flash overseas pricing. The compatibility names deepseek-chat and deepseek-reasoner are scheduled for deprecation on 2026-07-24 15:59 UTC. | DeepSeek models & pricing |
| DeepSeek | deepseek-v4-pro | $0.435 (cache miss) | $0.003625 (cache hit) | $0.87 | Official DeepSeek V4 Pro overseas pricing as listed on the current models and pricing page. | DeepSeek models & pricing |
| Qwen / Alibaba Cloud | qwen3.7-max | $2.50 (International, up to 1M) | $0.25 (explicit cache hit) | $7.50 (International, up to 1M) | Current Qwen-Max mainline model in Alibaba Cloud Model Studio. Explicit cache hits cost 10% of standard input; implicit cache hits cost 20%. The older qwen3-max family is scheduled for retirement on October 10, 2026. | Alibaba Cloud Model Studio pricing |
| Together AI / Z AI | GLM-5.2 | $1.40 | $0.26 | $4.40 | Together AI's public serverless price for Z AI's GLM-5.2, released June 16, 2026 with a 256K context window. | Together AI GLM-5.2 pricing |
| Meta / Llama via Together AI | Llama 4 Maverick | $0.27 | - | $0.85 | Meta's Llama 4 Maverick served through Together AI's public serverless API. Meta's public developer docs describe the model family, while Together AI exposes a comparable public token price. | Together AI Llama 4 Maverick pricing |
| Moonshot AI / Kimi | kimi-k2.7-code | $0.95 (cache miss) | $0.19 (cache hit) | $4.00 | Kimi K2.7 Code is Kimi’s current most capable coding model. The high-speed variant is $1.90 cache-miss input, $0.38 cache-hit input, and $8.00 output per 1M tokens. | Kimi K2.7 Code pricing |
| Xiaomi / MiMo | mimo-v2.5-pro | $0.435 (overseas cache miss) | $0.0036 (overseas cache hit) | $0.87 (overseas) | Current overseas pay-as-you-go price for Xiaomi MiMo V2.5 Pro. The V2 compatibility models were fully deprecated on June 30, 2026; V2.5 Pro supports a 1M context window and up to 128K output. | Xiaomi MiMo pay-as-you-go pricing |
| MiniMax | MiniMax-M3 | ¥2.10 (<=512K, 50% off) | ¥0.42 (cache read, <=512K) | ¥8.40 (<=512K, 50% off) | Current MiniMax pay-as-you-go text pricing for MiniMax-M3 standard service after the listed permanent 50% discount. Prompts above 512K are listed at ¥4.20 input, ¥0.84 cache read, and ¥16.80 output while limited supply applies; priority service is 1.5x standard pricing. | MiniMax pay-as-you-go pricing |
| xAI | Grok 4.5 | $2.00 | $0.50 | $6.00 | Current xAI flagship model with a 500K-token context window. Standard API pricing applies to text input, cached input, and text output. | xAI Grok 4.5 model documentation |
| GitHub Copilot | Copilot product pricing | N/A | N/A | N/A | GitHub Copilot is sold as a subscription product. GitHub does not publicly publish a Copilot per-token API price comparable to the other vendors here. | GitHub Copilot plans |
Subscription Plans
Publicly listed consumer or team subscriptions from the official provider pages I checked for this tracker.
| Vendor | Plan | Price | Notes | Official |
|---|---|---|---|---|
| OpenAI | ChatGPT Plus | $20 / month | Individual plan. API usage is billed separately. | OpenAI ChatGPT Plus help |
| OpenAI | ChatGPT Pro 5x | $100 / month | Pro capabilities with 5x the Plus usage allowance. No annual billing. | OpenAI ChatGPT Pro tiers |
| OpenAI | ChatGPT Pro 20x | $200 / month | Highest individual ChatGPT tier, with 20x the Plus usage allowance. No annual billing. | OpenAI ChatGPT Pro tiers |
| OpenAI | ChatGPT Business | $20 / user / month annually or $25 monthly | Standard ChatGPT seat for teams of two or more; pricing was reduced in April 2026. Flexible credits can extend included usage. | OpenAI ChatGPT Business help |
| Anthropic | Claude Pro | $20 / month or $200 / year | Individual plan; annual billing is equivalent to about $17 per month. | Claude plans and pricing |
| Anthropic | Claude Max 5x | $100 / month | Five times the Pro usage allowance. Includes Claude Code and Claude Cowork. | Claude plans and pricing |
| Anthropic | Claude Max 20x | $200 / month | Twenty times the Pro usage allowance. Includes Claude Code and Claude Cowork. | Claude plans and pricing |
| Anthropic | Claude Team Standard | $20 / seat / month annually or $25 monthly | For teams of 5 to 150. Anthropic also lists Premium seats at $100 annually or $125 monthly per seat. | Claude plans and pricing |
| Google AI Plus | $4.99 / month (US) | Individual plan with 2x Gemini usage limits and 400 GB of storage. Regional prices vary. | Google AI plans | |
| Google AI Pro | $19.99 / month (US) | Individual plan with 4x Gemini usage limits and 5 TB of storage. Regional prices vary. | Google AI plans | |
| Google AI Ultra 5x | $99.99 / month (US) | Higher-access Ultra tier with 5x the Pro usage limits and 20 TB of storage. | Google AI plans | |
| Google AI Ultra 20x | $199.99 / month (US) | Highest-access Ultra tier with 20x the Pro usage limits and 30 TB of storage. | Google AI plans | |
| Qwen / Alibaba Cloud | Model Studio Coding Plan Pro | $50 / month | Fixed-price coding plan for supported Qwen and selected third-party models; request quotas apply. | Alibaba Cloud Coding Plan |
| Xiaomi / MiMo | Token Plan (monthly) | $6 / $16 / $50 / $100 per month | Lite, Standard, Pro, and Max tiers covering the MiMo V2.5 model family. | Xiaomi MiMo Token Plan |
| Xiaomi / MiMo | Token Plan (annual) | $63.36 / $168.96 / $528 / $1,056 per year | Annual Lite, Standard, Pro, and Max tiers. | Xiaomi MiMo Token Plan |
| MiniMax | Token Plan Plus | ¥49 / month | Personal Token Plan for light development and daily use. | MiniMax Token Plan |
| MiniMax | Token Plan Max | ¥119 / month | Personal Token Plan for frequent coding agents and multimodal calls. | MiniMax Token Plan |
| MiniMax | Token Plan Ultra | ¥469 / month | Highest personal Token Plan for intensive agent workflows. | MiniMax Token Plan |
| xAI / Grok | X Premium+ | $40 / month or $395 / year (US web) | Includes higher Grok limits and SuperGrok; regional prices vary. | X Premium pricing |
| GitHub Copilot | Copilot Pro | $10 / user / month | Individual plan with $15 in monthly GitHub AI Credits. GitHub is gradually enabling new sign-ups. | GitHub Copilot plans |
| GitHub Copilot | Copilot Pro+ | $39 / user / month | Individual plan with $70 in monthly GitHub AI Credits. GitHub is gradually enabling new sign-ups. | GitHub Copilot plans |
| GitHub Copilot | Copilot Max | $100 / user / month | Highest individual tier with $200 in monthly GitHub AI Credits. GitHub is gradually enabling new sign-ups. | GitHub Copilot plans |
| GitHub Copilot | Copilot Business | $19 / seat / month | Organization-managed plan. Availability for new self-serve sign-ups can be staged by GitHub. | GitHub Copilot plan documentation |
| GitHub Copilot | Copilot Enterprise | $39 / seat / month | Enterprise-managed plan with organization-wide features and controls. | GitHub Copilot plan documentation |
DeepSeek and Moonshot AI / Kimi are omitted here because I could not find a public official consumer or coding subscription with a directly listed price in the sources checked for this snapshot.
Price History
Switch provider, metric, and time grain to compare the official pricing checkpoints I have stored so far.
History lines use official pricing snapshots that are stored in this repo. Some providers only have one official snapshot recorded so far, while others mix successive flagship or promo models.
Artificial Analysis Benchmark Snapshot
The current top 10 models on the Artificial Analysis Intelligence Index. This list is intentionally model-level, so repeated vendors can appear more than once when they occupy multiple top-10 slots.
Top 10 current models by Artificial Analysis Intelligence Index. Speeds and available prices come from the benchmark snapshot; missing model-level prices are filled from the official provider prices reviewed above. Values can differ when multiple deployments or reasoning modes exist. Source: Artificial Analysis models leaderboard (checked July 13, 2026).
| Vendor | Benchmark model | Intelligence | Speed | Blended price | Prompt pricing | Details |
|---|---|---|---|---|---|---|
| Anthropic | Claude Fable 5 (with fallback) | 59.86 | 60.37 t/s | $20 | Input $10 · Output $50 | Model details |
| OpenAI | GPT-5.6 Sol (max) | 58.89 | 69.05 t/s | $11.25 | Input $5 · Output $30 | Model details |
| Anthropic | Claude Opus 4.8 (max) | 55.69 | 58.48 t/s | $10 | Input $5 · Output $25 | Model details |
| OpenAI | GPT-5.6 Terra (max) | 54.95 | 144.38 t/s | $5.625 | Input $2.5 · Output $15 | Model details |
| OpenAI | GPT-5.5 (xhigh) | 54.84 | 77.55 t/s | $11.25 | Input $5 · Output $30 | Model details |
| xAI | Grok 4.5 (high) | 53.83 | 119.27 t/s | $3 | Input $2 · Output $6 | Model details |
| Anthropic | Claude Sonnet 5 (max) | 53.35 | 78.47 t/s | $4 | Input $2 · Output $10 | Model details |
| OpenAI | GPT-5.6 Luna (max) | 51.24 | 225.63 t/s | $2.25 | Input $1 · Output $6 | Model details |
| Z AI | GLM-5.2 (max) | 51.09 | 206.11 t/s | $2.15 | Input $1.4 · Output $4.4 | Model details |
| Muse | Muse Spark 1.1 (xhigh) | 50.62 | 121.37 t/s | $2 | Input $1.25 · Output $4.25 | Model details |
Top-25 API Provider Leaderboard
Artificial Analysis provider leaderboard snapshot, aggregated to the best currently benchmarked endpoint for each provider. Switch metrics to compare intelligence, blended price, latency, and context window separately.
Each provider rank uses that provider's best currently benchmarked endpoint for the selected metric, so the representative model can change across metrics. Source: Artificial Analysis provider leaderboard (checked July 13, 2026).
Scale and Price Frontier
A Star History-style yearly line for open-weight model sizes and the highest reviewed output-token prices since 2021.
Lines start in 2021 and combine curated source-backed historical records with the latest Artificial Analysis model metadata. Model size tracks the largest open-weight/open-access LLM by total disclosed parameters for each year, so sparse MoE and dense models are not quality-equivalent. Output-token price uses public USD text-token API prices and excludes tool-call, image, audio, video, and subscription pricing. Source: Artificial Analysis models (checked July 13, 2026).