LLM Pricing Tracker: API and Subscription Costs
A tracker for leading LLM API token prices and consumer subscriptions, with official links, repo snapshot refreshes, price-history charts, and live benchmark/provider snapshots.
Note
Latest repo snapshot in this build: September 25, 2026. Click Check for newer snapshot below to query GitHub for a fresher snapshot. Each browser is limited to one remote check per day.
This page tracks public pricing from official provider pages for major model vendors I regularly compare: OpenAI (GPT-6 Astra, Sol, and Luna), Google (Gemini 3.8 Flash and Gemma 4), Anthropic (Claude Fable 5.1, Opus 5.5, and Sonnet 5), xAI (Grok 4.7), DeepSeek (V4.1 Flash and V4 Pro), Qwen3.8 Max, Moonshot/Kimi (Kimi K3 and K2.7 Code), Xiaomi/MiMo V2.6, MiniMax M3, Together AI’s GLM-5.3 endpoint, and GitHub Copilot. Claude Mythos 5.1 is listed separately as invitation-only.
API and subscription prices were manually reviewed on September 25, 2026. Benchmark refresh dates below are separate from this pricing review date.
A few quick cautions before using the numbers:
- API pricing and consumer subscription pricing are different products.
- Some vendors publish tiered pricing by context length, region, or prompt type.
- API rows use standard-service USD prices per one million text tokens unless explicitly marked otherwise. Cache reads, cache writes, storage, tool calls, taxes, and subscription allowances are different charges.
- DeepSeek chart values use peak prices; off-peak rates are half as much. Gemini 3.8 Flash’s listed promotion ends on December 31, 2026. Qwen3.8 Max cache-hit prices require checking the Model Studio console and are not inferred from older models’ discounts.
- When a vendor does not publicly expose a comparable token-billing number, I mark that clearly instead of guessing.
- The charts below use snapshots stored in this repo, including daily-refreshable Artificial Analysis benchmark/provider snapshots and manually curated pricing rows from official vendor pages.
- Historical points retain the price observed on that date, including earlier models and promotions. Benchmark scores can change when Artificial Analysis revises its evaluation suite; they are not a fixed year-to-year scale.
API Token Fees
Representative API pricing for major frontier-model vendors. Numbers are taken from official vendor pricing pages; mixed currencies are kept in the vendor’s quoted currency.
| Vendor | Model / Product | Input | Cached input | Output | Notes | Official |
|---|---|---|---|---|---|---|
| OpenAI | GPT-6 Astra (<=272K input) | $10.00 | $1.00 | $50.00 | Standard service. Cache writes cost $12.50 per 1M tokens. Above 272K input tokens, the full request costs 2x input/cache rates and 1.5x output rates. Regional processing adds 10%; Batch/Flex and Fast have separate rates. | OpenAI API pricing |
| OpenAI | GPT-6 Sol (<=272K input) | $2.00 | $0.20 | $10.00 | Standard service. Cache writes cost $2.50 per 1M tokens. Above 272K input tokens, the full request costs 2x input/cache rates and 1.5x output rates. Regional processing adds 10%. | OpenAI API pricing |
| OpenAI | GPT-6 Luna (<=272K input) | $0.10 | $0.01 | $0.50 | Standard service. Cache writes cost $0.125 per 1M tokens. Above 272K input tokens, the full request costs 2x input/cache rates and 1.5x output rates. Regional processing adds 10%. | OpenAI API pricing |
| Gemini 3.8 Flash | $0.75 (through 2026-12-31) | $0.075 (through 2026-12-31) | $3.75 (through 2026-12-31) | Current promotional standard-tier pricing for Google's most capable Flash model. On January 1, 2027, prices become $1.50 input, $0.15 cached input, and $7.50 output; cache storage also rises from $0.50 to $1.00 per 1M tokens per hour. | Gemini API pricing | |
| Gemma 4 | Free of charge | Free of charge | Free of charge | Google's current Gemini Developer API pricing page lists Gemma 4 with free input, output, and context caching, while the paid tier remains unavailable. | Google Gemini Developer API pricing | |
| Anthropic | Claude Fable 5.1 | $10.00 | $0.25 (cache read) | $50.00 | Public API model. Cache reads cost 2.5% of base input, not the older 10% rate. Cache writes cost $12.50 (5 minutes) or $20 (1 hour) per 1M tokens. US-only inference is 1.1x standard pricing. | Claude API pricing |
| Anthropic | Claude Mythos 5.1 (limited access) | $10.00 | $0.25 (cache read) | $50.00 | Invitation-only Project Glasswing model; not a self-serve offering. Shares Fable 5.1 specifications and pricing. Cache writes cost $12.50 (5 minutes) or $20 (1 hour) per 1M tokens. | Anthropic Mythos 5.1 availability and pricing |
| Anthropic | Claude Opus 5.5 | $4.00 | $0.20 (cache read) | $20.00 | Public API model. Cache reads cost 5% of base input. Cache writes cost $5 (5 minutes) or $8 (1 hour) per 1M tokens. US-only inference is 1.1x standard pricing. | Claude API pricing |
| Anthropic | Claude Sonnet 5 | $2.00 | $0.20 (cache read) | $10.00 | Published rates rechecked September 25, 2026. Cache writes cost $2.50 (5 minutes) or $4 (1 hour) per 1M tokens. | Claude API pricing |
| DeepSeek | DeepSeek V4.1 Flash (deepseek-flash) | $0.30 peak / $0.15 off-peak | $0.006 peak / $0.003 off-peak | $1.20 peak / $0.60 off-peak | Use API ID deepseek-flash; retired V4 Flash aliases now route here. Supports vision, 1M context and up to 384K output. Peak hours: 01:00-04:00 and 06:00-10:00 UTC, Monday-Friday excluding Chinese public holidays. All other times are half-price; charts use peak rates. | DeepSeek models & pricing |
| DeepSeek | DeepSeek V4 Pro 0813 | $1.32 peak / $0.66 off-peak | $0.044 peak / $0.022 off-peak | $3.96 peak / $1.98 off-peak | API ID deepseek-v4-pro; V4-Pro-0813, 1M context and up to 384K output; no vision support. Peak hours: 01:00-04:00 and 06:00-10:00 UTC, Monday-Friday excluding Chinese public holidays. All other times are half-price; charts use peak rates. | DeepSeek models & pricing |
| Qwen / Alibaba Cloud | qwen3.8-max | $2.00 (Singapore / International) | See Model Studio console | $6.00 (Singapore / International) | Pay-as-you-go for up to 1M input tokens, both thinking and non-thinking. Other deployment regions have different rates. The cache documentation explicitly excludes Qwen3.8 Max from the generic 10% explicit / 20% implicit cache-hit rules and directs users to the console; no cache price is assumed here. | Alibaba Cloud Model Studio pricing |
| Together AI / Z AI | GLM-5.3 | $1.40 | $0.26 | $4.40 | Together AI's public serverless price for Z AI's GLM-5.3, released August 13, 2026 with a 1M-token context window. These are Together's rates, not Z AI's direct API rates. | Together AI GLM-5.3 pricing |
| Moonshot AI / Kimi | kimi-k3 | $3.00 (cache miss) | $0.30 (cache hit) | $15.00 | Open-weight native multimodal model with a 1,048,576-token context window. Cache writes cost $3 (5-minute TTL) or $6 (1-hour TTL) per 1M tokens, separate from $0.30 cache reads. | Kimi API Platform pricing |
| Moonshot AI / Kimi | kimi-k2.7-code | $0.95 (cache miss) | $0.19 (cache hit) | $4.00 | Dedicated coding model with a 262,144-token context window. The high-speed variant is $1.90 cache-miss input, $0.38 cache-hit input, and $8.00 output per 1M tokens. | Kimi K2.7 Code pricing |
| Xiaomi / MiMo | mimo-v2.6-pro | $0.435 (overseas cache miss) | $0.0036 (overseas cache hit) | $0.87 (overseas) | Overseas real-time API pricing; cache writes are temporarily free. Batch is half-price. The separate UltraSpeed variant costs $4.35 input, $0.036 cached input and $8.70 output. MiMo V2.5 Pro and V2.5 retire October 21, 2026 at 10:00 Beijing time. | Xiaomi MiMo pay-as-you-go pricing |
| Xiaomi / MiMo | mimo-v2.6-flash | $0.14 (overseas cache miss) | $0.0028 (overseas cache hit) | $0.28 (overseas) | Overseas real-time API pricing. Batch costs $0.07 input, $0.0014 cached input and $0.14 output per 1M tokens. Cache writes are temporarily free; web search is billed separately. | Xiaomi MiMo pay-as-you-go pricing |
| MiniMax | MiniMax-M3 | $0.30 (<=512K) | $0.06 (cache read, <=512K) | $1.20 (<=512K) | Global standard-service rates after the published permanent 50% reduction. Above 512K input tokens: $0.60 input, $0.12 cache read and $2.40 output. Priority service costs 1.5x standard rates. | MiniMax global API pricing |
| xAI | Grok 4.7 (<200K input) | $2.00 | $0.50 | $6.00 | 500K context window, but the long-context billing threshold is 200K input tokens. At or above 200K, all tokens in the request cost $4 input, $1 cached input and $12 output per 1M tokens. US regional inference adds 10%. | xAI API pricing |
| GitHub Copilot | Copilot product pricing | N/A | N/A | N/A | Copilot combines seat subscriptions with model-dependent GitHub AI Credit usage. There is no single Copilot-model input/output token rate; compare plan allowances below, not N/A values as zero-cost inference. | GitHub Copilot plans |
Subscription Plans
Publicly listed consumer or team subscriptions from the official provider pages I checked for this tracker.
| Vendor | Plan | Price | Notes | Official |
|---|---|---|---|---|
| OpenAI | ChatGPT Plus | $20 / month | Individual plan. API usage is billed separately. | OpenAI ChatGPT Plus help |
| OpenAI | ChatGPT Pro 5x | $100 / month | Pro capabilities with 5x the Plus usage allowance. No annual billing. | OpenAI ChatGPT Pro tiers |
| OpenAI | ChatGPT Pro 20x | $200 / month | 20x the Plus usage allowance. New sign-ups and upgrades paused since September 10, 2026; existing subscriptions continue, with a limited one-time return option for eligible former subscribers. No annual billing. | OpenAI ChatGPT Pro tiers |
| OpenAI | ChatGPT Business Standard | $20 / user / month annually or $25 monthly | Standard seat includes ChatGPT, ChatGPT Work and Codex. At least two Standard/Premium seats combined are required. API usage is separate; workspace credits can extend usage. | OpenAI ChatGPT Business help |
| OpenAI | ChatGPT Business Premium | $100 / user / month annually or $125 monthly | 5x Standard-seat usage and no 5-hour usage limit. Standard and Premium seats can be mixed; at least two seats combined are required. API usage is separate. | OpenAI ChatGPT Business help |
| Anthropic | Claude Pro | $20 / month or $200 / year | Individual plan; annual billing is equivalent to about $17 per month. | Claude plans and pricing |
| Anthropic | Claude Max 5x | $100 / month | Five times the Pro usage allowance. Includes Claude Code; model-specific limits apply. | Claude plans and pricing |
| Anthropic | Claude Max 20x | $200 / month | Twenty times the Pro usage allowance. Includes Claude Code; model-specific limits apply. | Claude plans and pricing |
| Anthropic | Claude Team Standard | $20 / seat / month annually or $25 monthly | For teams of 2 to 150. Premium seats cost $100 per seat per month billed annually, or $125 per seat per month billed monthly; seat types can be mixed. | Claude plans and pricing |
| Google AI Plus | $4.99 / month (US) | Individual plan with 2x Gemini usage limits and 400 GB of storage. Regional prices and storage bundles vary. | Google AI plans | |
| Google AI Pro | $19.99 / month (US) | Individual plan with 4x Gemini usage limits and 5 TB of storage. Regional prices vary. | Google AI plans | |
| Google AI Ultra 5x | $99.99 / month (US) | Higher-access Ultra tier with 5x the Pro usage limits and 20 TB of storage. | Google AI plans | |
| Google AI Ultra 20x | $199.99 / month (US) | Highest-access Ultra tier with 20x the Pro usage limits and 30 TB of storage. | Google AI plans | |
| Qwen / Alibaba Cloud | Model Studio Coding Plan Pro | $50 / month | Fixed-price coding plan for supported Qwen and selected third-party models; request quotas apply. | Alibaba Cloud Coding Plan |
| Qwen / Alibaba Cloud | Model Studio Token Plan (Individual) | $8 / $25 / $80 per month (list) | Lite, Standard and Pro list prices. The launch announcement advertises limited-time $6 / $18 / $68 rates; confirm eligibility and current discounts at checkout. Credits have 5-hour and 7-day windows, separate from pay-as-you-go API billing. | Alibaba Cloud Model Studio Token Plan |
| Xiaomi / MiMo | Token Plan (monthly) | $6 / $16 / $50 / $100 per month | Individual Lite, Standard, Pro and Max developer tiers; includes MiMo V2.6 Pro/Flash. Separate coding-tool quota, not general-purpose API balance or a MiMo Desktop membership price. | Xiaomi MiMo Token Plan |
| Xiaomi / MiMo | Token Plan (annual) | $63.36 / $168.96 / $528 / $1,056 per year | Individual Lite, Standard, Pro and Max annual developer plans, billed upfront. Rates reflect the documented annual discount; separate from Desktop membership pricing. | Xiaomi MiMo Token Plan |
| MiniMax | Token Plan Plus | $22 / month | Personal Token Plan with rolling 5-hour and weekly quota windows. Optional prepaid Credits cover overflow and are billed separately. | MiniMax global Token Plan |
| MiniMax | Token Plan Max | $55 / month | Personal Token Plan with rolling 5-hour and weekly quota windows. Text, eligible image and speech resources share the quota. | MiniMax global Token Plan |
| MiniMax | Token Plan Ultra | $132 / month | Personal Token Plan for intensive use, with rolling 5-hour and weekly quota windows. Not an unlimited API subscription. | MiniMax global Token Plan |
| Moonshot AI / Kimi | Kimi Plus | $19 / month or $180 / year | USD prices on the Kimi Code plan page; annual equivalent $15/month. Includes Kimi membership and coding access. New plans have monthly quotas and a rolling 5-hour limit; legacy plans have different rules. | Kimi Code membership plans |
| Moonshot AI / Kimi | Kimi Pro | $39 / month or $372 / year | USD prices; annual equivalent $31/month. Kimi Code with K3 up to 1M context and HighSpeed access. API Open Platform billing is separate. | Kimi Code membership plans |
| Moonshot AI / Kimi | Kimi Max | $99 / month or $948 / year | USD prices; annual equivalent $79/month. Higher membership quota, shared with Kimi Code; API Open Platform billing is separate. | Kimi Code membership plans |
| Moonshot AI / Kimi | Kimi Ultra | $199 / month or $1,908 / year | USD prices; annual equivalent $159/month. Highest coding-plan quota shown on the public page; API Open Platform billing is separate. | Kimi Code membership plans |
| xAI / Grok | X Premium+ | $40 / month or $395 / year (US web) | Includes higher Grok limits and SuperGrok; regional prices vary. | X Premium pricing |
| GitHub Copilot | Copilot Pro | $10 / user / month | Currently $15 monthly AI Credit value: $10 base plus $5 variable Flex allotment. Flex allowances can change. | GitHub Copilot plans |
| GitHub Copilot | Copilot Pro+ | $39 / user / month | Currently $70 monthly AI Credit value: $39 base plus $31 variable Flex allotment. Flex allowances can change. | GitHub Copilot plans |
| GitHub Copilot | Copilot Max | $100 / user / month | Currently $200 monthly AI Credit value: $100 base plus $100 variable Flex allotment. Flex allowances can change. | GitHub Copilot plans |
| GitHub Copilot | Copilot Business | $19 / seat / month | Organization-managed plan; 1,900 AI Credits ($19 value) per granted seat per month contribute to a shared pool. Additional usage is billed separately. | GitHub Copilot plan documentation |
| GitHub Copilot | Copilot Enterprise | $39 / seat / month | For GitHub Enterprise Cloud; 3,900 AI Credits ($39 value) per granted seat per month contribute to a shared pool. Additional usage is billed separately. | GitHub Copilot plan documentation |
DeepSeek has no comparable paid subscription price verified in this review. Subscription amounts are USD unless labeled otherwise; regional taxes, checkout offers and legacy-plan eligibility can differ. API billing is separate from consumer and coding subscriptions.
Price History
Switch provider, metric, and time grain to compare the official pricing checkpoints I have stored so far.
History lines use official pricing snapshots that are stored in this repo. Some providers only have one official snapshot recorded so far, while others mix successive flagship or promo models.
Artificial Analysis Benchmark Snapshot
The current top 10 models on the Artificial Analysis Intelligence Index. This list is intentionally model-level, so repeated vendors can appear more than once when they occupy multiple top-10 slots.
Top 10 non-deprecated model/reasoning configurations by Artificial Analysis Intelligence Index. Scores, speeds and prices are from the same dated snapshot; benchmark versions can change, so scores are not necessarily comparable with older snapshots. Missing values remain unavailable. Blended prices use uncached input/output tokens in a 3:1 ratio, not AA's cache-weighted mix. Source: Artificial Analysis models leaderboard (checked September 25, 2026).
| Vendor | Benchmark model | Intelligence | Speed | Blended price | Prompt pricing | Details |
|---|---|---|---|---|---|---|
| Anthropic | Claude Opus 5.5 (max with fallback) | 57.62 | t/s | $8 | Input $4 · Output $20 | Model details |
| Anthropic | Claude Opus 5.5 (xhigh with fallback) | 55.99 | 91.46 t/s | $8 | Input $4 · Output $20 | Model details |
| Anthropic | Claude Opus 5.5 (high with fallback) | 53.58 | 87.63 t/s | $8 | Input $4 · Output $20 | Model details |
| Anthropic | Claude Fable 5.1 (max with fallback) | 53.35 | 69.07 t/s | $20 | Input $10 · Output $50 | Model details |
| Anthropic | Claude Fable 5.1 (xhigh with fallback) | 53.2 | 60.39 t/s | $20 | Input $10 · Output $50 | Model details |
| OpenAI | GPT-6 Astra (max) | 52.67 | 53.72 t/s | $20 | Input $10 · Output $50 | Model details |
| OpenAI | GPT-6 Astra (xhigh) | 52.39 | 52.66 t/s | $20 | Input $10 · Output $50 | Model details |
| Anthropic | Claude Opus 5.5 (medium with fallback) | 51.24 | 79.13 t/s | $8 | Input $4 · Output $20 | Model details |
| Anthropic | Claude Fable 5.1 (high with fallback) | 51.15 | 56.91 t/s | $20 | Input $10 · Output $50 | Model details |
| OpenAI | GPT-6 Astra (high) | 50.92 | 53.07 t/s | $20 | Input $10 · Output $50 | Model details |
Top-25 API Provider Leaderboard
Artificial Analysis provider leaderboard snapshot, aggregated to the best currently benchmarked endpoint for each provider. Switch metrics to compare intelligence, blended price, latency, and context window separately.
Each provider rank uses that provider's best currently benchmarked endpoint for the selected metric, so the representative model can change across metrics. Source: Artificial Analysis provider leaderboard (checked September 25, 2026).
Scale and Price Frontier
A Star History-style yearly line for open-weight model sizes and the highest reviewed output-token prices since 2021.
Lines start in 2021 and combine curated source-backed historical records with the latest Artificial Analysis model metadata. Model size tracks the largest open-weight/open-access LLM by total disclosed parameters for each year, so sparse MoE and dense models are not quality-equivalent. Output-token price uses public USD text-token API prices and excludes tool-call, image, audio, video, and subscription pricing. Source: Artificial Analysis models (checked September 25, 2026).