Quick Answer
Claude API pricing 2026: Fable 5.1 $10 input / $50 output per million tokens with $0.25 cache reads (0.025x — 4x cheaper than the standard 0.1x); Sonnet 5 $2/$10 PERMANENT (intro pricing made permanent August 10, 2026); Opus 5 $5/$25; Haiku 4.5 $1/$5. Batch API 50% off; Bedrock and Vertex AI add 10% regional premium on certain endpoints.
Last verified: Sep 16, 2026.
At a glance
- Claude Fable 5.1: $10 input / $50 output per 1M tokens; cache read $0.25 (0.025x)
- Claude Mythos 5.1: $10/$50 same as Fable; Glasswing-restricted
- Claude Sonnet 5: $2/$10 PERMANENT; default for Free and Pro plans
- Claude Opus 5 / Opus 4.8: $5 input / $25 output per 1M tokens
- Claude Haiku 4.5: $1 input / $5 output; cheapest in the family
- Cache writes: $12.50 (5-min) / $20 (1-hour) for Fable 5.1
- Batch API: 50% off standard; 24-hour turnaround
- Regional uplift: 10% premium on Bedrock/Vertex regional endpoints; 1.1x US-only inference multiplier
Fable 5.1 cache reads are the cheapest in frontier AI
Claude Fable 5.1's $0.25-per-million-token cache read is 4x cheaper than OpenAI's cache read on GPT-6 Astra, and it is the single biggest cost lever for high-volume Claude deployments.
Anthropic uses two cache multipliers: 0.025x on Fable 5.1 and Mythos 5.1, and 0.1x on every other model including Opus 5, Sonnet 5, and Haiku 4.5. That means a Fable 5.1 cache hit is $0.25 per million tokens versus $1.00 on Fable 5, $0.50 on Opus 5, $0.20 on Sonnet 5, and $0.10 on Haiku 4.5. For a workload that reuses a 50,000-token system prompt across 10,000 requests per day, Fable 5.1 saves $375 per day versus Sonnet 5 on cache reads alone (Anthropic pricing page, September 2026).
The cache is automatic: any prompt prefix that has been seen in the last 5 minutes (or 1 hour if configured) is served from cache at the cache read rate. Cache writes happen once per prefix per TTL window, so the amortised cost depends on reuse rate. A 50,000-token prefix written once and reused 100 times in 5 minutes costs $0.0125 per request on Fable 5.1 — 800x cheaper than the headline input rate (eesel.ai, September 2026).
Sonnet 5 became the default in mid-2026
Claude Sonnet 5 became the default Claude model for Free and Pro plans in Claude.ai on June 30, 2026, and Anthropic made the introductory $2/$10 pricing permanent on August 10, 2026.
The launch pricing was originally set to expire August 31, 2026, with the standard rate scheduled to rise to $3 input / $15 output on September 1, 2026. Anthropic cancelled that increase in an August 10 edit to the Sonnet 5 launch post, citing Sonnet 5's competitive position and the cost benefit to power users (Anthropic blog, August 2026).
Sonnet 5 has a 1M-token context window and 128K max output — matching Fable 5.1 and Opus 5. For most workloads that do not need Fable 5.1's deeper reasoning, Sonnet 5 offers the best price-performance in the Claude family: $2/$10 is 5x cheaper than Fable 5.1 on input and 2.5x cheaper on Opus 5. Sonnet 4.6 and 4.5 remain available at $3/$15 for teams that prefer the older model family (Anthropic pricing page, September 2026).
Opus 5 vs Opus 4.x: same price, newer knowledge cutoff
Anthropic priced Opus 5 identically to Opus 4.8, 4.7, 4.6, and 4.5 — the choice is about capabilities, not cost.
All five Opus models list at $5 input / $25 output per million tokens. Cache writes are $6.25 (5-minute) / $10 (1-hour) and cache reads are $0.50 per million tokens on all five. Fast mode on Opus 5 and Opus 4.8 doubles the input and output rates to $10/$50 per million tokens. The Opus 4.x line is the cheapest path to a 1M-token Claude context window — Opus 5 does not have a price advantage over its predecessors (Anthropic pricing page, September 2026).
Opus 4.1 and Opus 4 are listed on Anthropic's pricing page but are retired on the first-party API. They remain available on AWS Bedrock and Google Cloud for legacy customers at the same $5/$25 rate. New Anthropic API customers cannot provision them. Migrating Opus 4.1 or 4.0 traffic to Opus 5 is a free path to a more recent knowledge cutoff (May 2026) at the same price (eesel.ai, September 2026).
What enterprise buyers should do next
Three actions for organisations evaluating Claude API pricing in 2026.
- Pick Sonnet 5 unless you have a reason not to. Sonnet 5's $2/$10 permanent rate is the best price-performance in the Claude family. Fable 5.1 is justified for deep reasoning and long-horizon agentic work. Opus 5 is justified only when Sonnet 5's accuracy is insufficient. Haiku 4.5 is the cheapest Claude API option for high-volume classification.
- Design for prompt caching from day one. Keep system prompts above 1,024 tokens and as stable across requests as possible. A 50K-token system prompt reused 100x in 5 minutes runs at $0.0125 per request on Fable 5.1 versus $0.50 per request without caching.
- Negotiate on annual commitments. Anthropic offers volume discounts for enterprise contracts starting around 100M tokens/month. The pricing page lists standard rates only — committed-use discounts can cut 20-40% off headline rates for predictable workloads.
What to watch next
Three near-term datapoints. First, whether Mythos 5.1 remains Glasswing-restricted — Anthropic has not yet published when (or if) Mythos 5.1 general availability will lift. Second, the Claude 5.2 release window — Anthropic typically ships a model family refresh every 6-9 months; a Sonnet 5.5 or Fable 5.5 is plausible for Q1 2027. Third, the impact of the 1.1x US-only inference multiplier on competitive pricing for US-regulated workloads — this multiplier applies on Fable 5.1, Opus 5, and Sonnet 5, making Anthropic US-only endpoints 10% more expensive than global endpoints.







