On September 3, 2026, OpenAI released GPT-6 Astra — a frontier model with a 1,050,000-token context window and the first model to reach Critical cybersecurity capability. Here is the full pricing breakdown, benchmark results, and how to access gpt-6-astra via API.
GPT-6 Astra — OpenAI's 1,050,000-token frontier model with Critical cybersecurity capability
GPT-6 Astra is OpenAI's latest frontier model, released on September 3, 2026. It features a 1,050,000-token context window — the largest of any OpenAI model to date — with a maximum output of 128,000 tokens. The model accepts text and images as input and produces text output, with a knowledge cutoff of April 30, 2026.
Astra is the first model to achieve the Critical cybersecurity capability level, meaning it can autonomously discover and exploit novel vulnerabilities in real-world systems. This makes it a significant step toward agentic security research, but also raises the bar for responsible deployment.
The model supports five reasoning effort levels — low, medium, high, xhigh, and max — with the API defaulting to low. This gives developers fine-grained control over the latency-cost tradeoff, from fast lightweight queries to deep multi-step reasoning.
GPT-6 Astra costs $10 per million input tokens and $50 per million output tokens at standard rates. Cached input is discounted to $1 per million tokens, and cache writes cost $12.50 per million tokens. Long prompts exceeding 272,000 input tokens trigger a 2x multiplier on input and cache rates, and a 1.5x multiplier on output for the full request.
| Model | Input (cache miss) | Input (cache hit) | Output | Context |
|---|---|---|---|---|
| GPT-6 Astra (gpt-6-astra) | $10.00 | $1.00 | $50.00 | 1,050,000 |
Source: OpenAI official pricing as of September 3, 2026. Long prompts (>272K input tokens) incur 2x input/cache and 1.5x output rates on the full request. Batch and Flex modes available at 50% of Standard rates. Fast mode delivers 2x speed at 2x price.
GPT-6 Astra supports three execution modes beyond Standard. Batch and Flex both run at 50% of Standard rates, making them cost-effective for workloads that do not need real-time responses. Fast mode doubles the speed at double the price, useful when latency matters more than cost.
| Mode | Input | Output | Notes |
|---|---|---|---|
| Standard | $10.00 | $50.00 | Default mode |
| Batch / Flex | $5.00 | $25.00 | 50% of Standard |
| Fast | $20.00 | $100.00 | 2x speed at 2x price |
Prompts exceeding 272,000 input tokens trigger a long-context surcharge. Input and cached input rates double, and output rates increase by 1.5x — applied to the entire request, not just the tokens above the threshold. For a 500K-token input prompt, you would pay $20/MTok input instead of $10, and $75/MTok output instead of $50.
This pricing structure rewards developers who keep prompts concise. If your workload consistently uses large prompts, batching related context into fewer, shorter calls can meaningfully reduce cost.
At $10/$50 per million tokens, GPT-6 Astra sits above GPT-5.6 Sol ($5/$30) and Kimi K3 ($3/$15) but below Claude Fable 5's premium tier. The 1,050K context window is the largest available from OpenAI, which can justify the higher per-token cost for workloads that genuinely need to process massive inputs in a single call.
Prompt caching has the largest effect on sustained agent sessions. At a 90% input-cache hit rate, a workload using 10 million input tokens and 1 million output tokens costs about $55 on Astra Standard, compared with $32 on Sol and $15.70 on K3. The economics shift when you factor in the larger context window — Astra can handle entire repositories in one call where smaller models need chunking and multiple passes.
GPT-6 Astra is available through the OpenAI API and on AWS Bedrock. It supports the OpenAI chat completions format, so existing integrations built for GPT-5.x models can switch by changing the model ID to gpt-6-astra. No new SDK is required.
Astra is best suited to workloads that need both massive context and strong reasoning: large-codebase analysis, multi-document research, and agentic security testing. The Critical cybersecurity capability level means it can autonomously find and exploit novel vulnerabilities — a capability that is both powerful and sensitive, requiring careful access controls.
The five reasoning effort levels give developers more flexibility than previous OpenAI models. Using low for straightforward tasks and reserving max for complex multi-step reasoning can keep costs manageable while maintaining quality where it matters.
APItokendeal provides OpenAI-compatible access to supported GPT, Claude, DeepSeek, Gemini, and Kimi models through one key. Teams can evaluate Astra against other model families without maintaining a separate integration for each provider.
APItokendeal provides OpenAI-compatible API access to a wide range of models including GPT-6 Astra, Claude, DeepSeek, Gemini, Kimi, and more. One API key, one dashboard, prepaid token packs with no surprise bills.