Last verified: October 5, 2026.
Most teams quote OpenAI pricing as one rate per model. The invoices tell a different story. Every GPT-6 Astra request runs through four separately billed meters — and then there is a fifth cost that never appears on a slide deck: a threshold that silently re-prices the entire request the moment a prompt gets long enough. The teams that ship profitably on this API are not the ones that know the headline rates. They are the ones that know which meter is actually running.
This piece dissects the full 2026 rate card — the four meters, the cliff, the cache-write math, the discount modes, and the model family tiers — with the worked numbers that show where budgets actually go.
Four Token Rates, Not One — and Only One Gets Quoted
OpenAI documents four billed categories on the API pricing page, and the distinction is not cosmetic. Uncached input, cached input, cache writes, and output each carry their own rate, and most teams do not realize cache writes fire automatically any time a prompt prefix runs longer than 1,024 tokens. Your system prompt is being billed twice — once to read, once to remember.
| Billed category | GPT-6 Astra rate | What triggers it |
|---|---|---|
| Uncached input | $10 / million tokens | Every fresh prompt token |
| Cached input | $1 / million tokens | Prefixes seen recently in the same project |
| Cache writes | $12.50 / million tokens | Any prefix over 1,024 tokens — automatic |
| Output | $50 / million tokens | Every generated token |
The cached-read discount is the steepest on the card: 90% off the uncached rate, served automatically for prefixes seen within the last 5 to 60 minutes depending on configuration. At the batch tier, cached reads drop further to $0.50 per million. The economics reward one design decision above all others — stable, reusable prompt prefixes — and punish the opposite.
The 272K Cliff That Re-Prices Your Entire Request
Here is the mechanism that turns a predictable bill into a surprise. When input tokens exceed 272,000, GPT-6 Astra does not bill just the overflowing portion at a premium. The entire request switches to the long-context rate — input doubles, output jumps 50%. A request at 273,000 tokens costs roughly twice what the same request costs at 271,000.
| Request state | Input rate | Output rate | Scope of the hike |
|---|---|---|---|
| 271,000 input tokens (short context) | $10 / million | $50 / million | Standard rates end to end |
| 273,000 input tokens (long context) | $20 / million | $75 / million | Entire request re-priced |
| Same doc split into two requests | $10 / million | $50 / million | Cliff never triggered |
The 272K figure is a pricing boundary, not a capacity boundary. The same model page lists a 1,050,000-token context window — up to 922,000 input and 128,000 output — so the model can technically ingest what the rate card punishes you for feeding it. That gap is why most production long-document RAG systems chunk and run separate requests rather than combine a full document into one submission. The cliff applies identically on Azure AI Foundry, AWS Bedrock when Astra ships there, and OpenRouter, so there is no provider detour around it.
Batch, Flex, Fast: Three Discounts That Aren't Equal
OpenAI sells the same frontier model at three different speeds, and the differences go deeper than price. All three apply the 50% discount logic across the catalogue, including the GPT-5.6 family — but the turnaround and the failure terms differ in ways that decide which workload belongs where.
| Mode | GPT-6 Astra rate | Turnaround | The catch |
|---|---|---|---|
| Batch | $5 in / $25 out | Within 24 hours | Useless for interactive flows |
| Flex | $5 in / $25 out | Within 2 hours | Looser SLAs, retry-on-failure free |
| Fast | $20 in / $100 out | Single-digit seconds | 2x standard — unavailable under EU residency |
For workloads that can wait, Batch is the cheapest route to a frontier model on the entire OpenAI catalogue — nightly RAG ingestion, bulk classification, offline evaluation. Flex exists for the awkward middle: interactive enough to need hours, tolerant enough to accept retries. Fast mode is the only path that beats Flex's two-hour window, and it bills accordingly.
The GPT-5.6 Family: Four Models, Four Different Jobs
Underneath the frontier sits a family ladder that most cost problems should climb before touching Astra at all. All four GPT-5.6 models inherit the same discount structure — the 50% Batch rate, the 2x Fast multiplier, the 1.25x cache-write premium — which makes the comparison clean:
| Model | Input / output per million | Slot in the lineup |
|---|---|---|
| GPT-5.6 Sol | $4 / $20 (promotional) | Mid-tier reasoning flagship; 2.5x cheaper than Astra |
| GPT-5.6 Terra | $2 / $12 | Cost-optimized general-purpose workhorse |
| GPT-5.6 Luna | $0.20 / $1.20 | Cheapest flagship-tier model — ~50x cheaper input than Astra |
| GPT-5.6 Cyber | $12.50 input | Safety-classified tier |
The Sol promotional rate holds at least through November 21, 2026, after the GPT-6 Astra launch on September 3, 2026, and OpenAI generally keeps prior models available for 90 days past a successor's debut. For routine workloads Sol already handles reliably, the Astra premium does not pay back. The 2.5x multiplier is justified only when the frontier model measurably cuts retries, completion time, or failed runs — which is a measurement question, not a faith question. One more rung below: GPT-5.3 Codex at $1.75 per million input tokens remains the cheapest coding-optimized option for generation workloads that need no reasoning at all.
Cache Writes: The Line Item That Eats Margins
Run the standard production scenario through the rate card and the hidden line item surfaces immediately. Take a 5,000-token system prompt plus a 20,000-token user prompt generating a 2,000-token answer. A single request incurs 25,000 uncached input tokens at the standard rate — $0.25. The 5,000-token prefix over the 1,024-token trigger bills a cache write at 1.25x the input rate — $0.0625. The answer adds $0.10 of output. Total: $0.4125 per request.
Scale that to 1,000 turns a day and the bill runs $412.50 daily — of which $62.50 is pure cache-write spend that most budgets never itemized. The defense is structural, not negotiable: prefixes that stay stable amortize the write once and read cheap forever after; prefixes that churn pay the 1.25x premium on every generation. Cache-read rates are generous precisely because OpenAI wants your prefix to persist. Design accordingly.
Regional Routing and the 10% You Didn't Budget
One more meter hides at the bottom of the rate card. Requests routed through regional processing endpoints — US, EU, or APAC data residency — carry a 10% uplift on GPT-6 Astra token rates. The base api.openai.com endpoint carries no surcharge, which means the compliance decision is also a line-item decision. And the interaction compounds: Fast mode is unavailable under EU data residency as of September 2026, so latency-critical EU workloads face an impossible trade — residency without speed, or speed without residency.
The Dates That Decide the Price
Three dates govern where this rate card goes next. The Sol promotional rate is guaranteed only through November 21, 2026 — after that, the mid-tier ladder re-prices on whatever the successor launch dictates. The 90-day availability rule means anything announced since Astra's September 3 debut could retire within the quarter. And the GPT-5 generation beneath it all — $1.25/$10 standard, $0.25/$2 mini, $0.05/$0.40 nano — rides the discount ladder downward as it ages, which is exactly where high-volume, low-complexity workloads belong.
None of these numbers sit still while you read them. The rate card above is the September 2026 snapshot; the cliff and the four meters, however, have survived every repricing so far — and those are the parts worth designing around.
Read next
Claude API Pricing 2026: Cost Per Million Tokens runs the same four-meter breakdown against Anthropic's card, and Gemini API Pricing 2026: Cost Per Million Tokens completes the three-way comparison your budget review actually needs.






