GPT-6 Astra API
GPT-6 Astra launched on 2026-09-03 as OpenAI’s most capable model. It lists at $10 input / $50 output per million tokens, and requests above 272K input tokens are billed at a higher rate in full. This page covers the official and third-party benchmarks, how it compares with Claude Opus 5.5, Fable 5.1 and GPT-6 Sol, how to set reasoning effort, its cybersecurity limits, community reviews, the parameters to remove when migrating, and Wokey pricing.
Price check: official, OpenRouter and Wokey
USD per million tokens. Official and OpenRouter rates were checked by hand on 2026-09-29, excluding tax; the Wokey rate is read from the live price list.
| Channel | Input | Cache read | Output |
|---|---|---|---|
| Official API | $10 | $1 | $50 |
| OpenRouter | $10 | $1 | $50 |
| Wokey | $0.9 | $0.09 | $4.5 |
GPT-6 Astra at a glance
- Release date
- 2026-09-03 for selected API organizations and Daybreak; from 09-04 ChatGPT Business, Pro and Plus followed, with Enterprise enabled by admins. Pro, Business and Enterprise also get Astra Pro
- Positioning
- OpenAI pitches it “for our highest level of capability”. GPT-6 Sol (released 09-22) targets demanding tasks and GPT-6 Luna targets scale, and OpenAI says Astra remains its best model across the board
- Context and output
- 1,050,000 context (922K max input) and 128K max output; text and image input, text output; knowledge cutoff 2026-04-30
- Reasoning effort
- low / medium / high / xhigh / max, with no none; configuration_update items change effort mid-conversation without breaking the cache prefix
- Official price
- $10 input / $50 output per million tokens, $1 cache read, $12.50 cache write; requests above 272K input are billed in full at 2x input and cache and 1.5x output ($20 / $2 / $25 / $75); Batch and Flex 50% off, Fast mode 2x
- Cybersecurity limits
- OpenAI rates its cyber capability as Critical under the Preparedness Framework. The public model refuses to write proof-of-concept exploits and is limited to secure code review and patching; full capability is offered through Daybreak
- Codex
- Codex adds context save and retrieval: notes carried across context windows, with earlier windows searchable. It is an experimental opt-in that will become the default; ChatGPT usage counts toward subscription allowances
- Model ID and platforms
- gpt-6-astra (the only snapshot); Chat Completions, Responses and Batch, with tool calling requiring Responses; also on AWS Bedrock and Microsoft Foundry; Tier 1 limits 500 RPM / 500K TPM
· Sources: OpenAI · Migration guide · Benchmarks (Anthropic)
Official benchmarks: GPT-6 Astra vs Opus 5.5 vs Fable 5.1 vs Opus 5
From Anthropic’s Opus 5.5 launch page, keeping only the benchmarks that list a GPT-6 Astra score. These are vendor-reported results, and GPT-6 Astra scores are OpenAI-reported.
| Benchmark | GPT-6 Astra | Opus 5.5 | Fable 5.1 | Opus 5 |
|---|---|---|---|---|
| Terminal-Bench 4.0 | 57.9% | 66.4% | 55.8% | 52.3% |
| FrontierCode v1.1 (Main) | 53.3% | 54.4% | 50.3% | 48.0% |
| GDPval-AA v2.1 | 1542 | 1846 | 1735 | 1708 |
| AutomationBench | 41.4% | 40.0% | 31.4% | 26.9% |
| Humanity's Last Exam (tools) | 57.2% | 67.7% | 65.6% | 63.6% |
| Terminal-Bench-Science 0.1 | 64.6% | 58.7% | 52.6% | 29.0% |
Anthropic compiled this table, and its benchmark selection leans toward Claude’s strengths. OpenAI’s launch reported a different set: 98% on FrontierMath Tier 4, 99.9% on ARC-AGI-3 and 100% on ExploitBench (GPT-5.6 Sol: 78.5%), which cannot be compared with the table directly.
Which to use: GPT-6 Astra, Claude Opus 5.5 or GPT-6 Sol
vs Claude Opus 5.5 / Fable 5.1
Opus 5.5 lists at $4 / $20, two-fifths of Astra’s price, and clearly leads on Terminal-Bench 4.0, GDPval-AA and Humanity’s Last Exam above; Astra leads on AutomationBench and Terminal-Bench-Science. On Artificial Analysis’ Intelligence Index Opus 5.5 scores 58 at max for $5.98 per task and Astra 53 for $3.26: Astra costs more per token but uses far fewer output tokens, so each task costs less. Fable 5.1 shares Astra’s $10 / $50 price and its score of 53, but costs $7.63 per task, more than twice as much.
vs GPT-6 Sol / Luna
GPT-6 Sol lists at $2 / $10, a fifth of Astra’s price, and GPT-6 Luna at $0.10 / $0.50. On Artificial Analysis’ Intelligence Index Sol scores 48 for $1.06 per task and Luna 37 for $0.07; on its Coding Agent Index Astra scores 62, Sol 57 and Luna 41. Start everyday coding and routine agent work on Sol and move to Astra for problems Sol cannot solve; use Luna for high-volume classification and extraction.
Choosing effort
Artificial Analysis measured low at 46 for $0.82 per task, medium at 50 for $1.54, high at 51 for $1.73 and max at 53 for $3.26, with every level on the cost-versus-score frontier. Start at medium: high costs about 12% more for 1 point, and max costs more than twice as much for 3 points. When migrating, replace none with low and start evaluating minimal at low.
Third-party and community reviews
- Artificial Analysis: It scores 53 on the Intelligence Index at max, tied for first with Fable 5.1 and 6 points above GPT-5.6 Sol; a task costs $3.26, about 40% of Fable 5.1’s $7.63 and about 60% more than GPT-5.6 Sol. It uses about 27K output tokens per task, a third of Fable 5.1’s 78K. It scores 62 on the Coding Agent Index, level with Fable 5.1 and ahead of Opus 5 (60) and GPT-5.6 Sol (55), at $7.09 per task, about 40% cheaper than Fable 5.1 and 30% cheaper than Opus 5. Its AA-Omniscience hallucination rate fell to 51% from GPT-5.6 Sol’s 92%. It trails GPT-5.6 Sol by about 45 Elo on GDPval-AA but needs only 24 turns per task, versus 45 for GPT-5.6 Sol and 60 for Fable 5.1 and Opus 5. Every Wokey meter is 9% of the official rate, so the same task comes to about $0.29 at max and $0.14 at medium.
- OpenAI system card and safety testing: In agentic tests without a confirmation policy, 3.4% of runs ended in misaligned outcomes versus 18.8% for GPT-5.6 Sol, and Gray Swan’s indirect prompt injection attacks succeeded 8.5% of the time versus 27.0%. A 100% ExploitBench score led OpenAI to rate a model’s cyber capability as Critical for the first time; the public model is limited to secure code review and patching, and Daybreak access will expand later. These figures are OpenAI-reported.
- Hacker News (2,279-point thread): An external Codex tester called it a “grounded collaborator and executor” that avoids doing work you did not ask for. Others found Fable better at inferring intent and said Astra needs very specific instructions, while some liked that it does not treat questions as instructions. After general availability, commenters reported usage at about 2.5x the rate of GPT-5.6 Sol, and one used up a 5-hour limit in 15 messages.
Migrating from GPT-5.6 Sol and older models: parameters to remove and behavior changes
These changes come from OpenAI’s model page and latest-model guide. Check your request parameters, cache settings, tool calling and prompts before switching the model to gpt-6-astra.
- Remove temperature, top_p and top_logprobs; on Chat Completions also remove logprobs, and on Responses drop message.output_text.logprobs from include. Astra owns its sampling policy and rejects requests with these fields. Through Wokey the gateway strips temperature, top_p, top_logprobs and logprobs before forwarding.
- reasoning.effort does not accept none: replace none with low, and start evaluating minimal at low. Through Wokey none and minimal are mapped to low automatically, so the request does not fail.
- Chat Completions can call Astra, but tool calling requires Responses, so move tool-using agents to /v1/responses. Through Wokey, Chat Completions requests are converted to Responses before forwarding, so Chat requests with tools also work.
- Caching: replace prompt_cache_retention with prompt_cache_options.ttl: "30m", and to change effort mid-session send a configuration_update item instead of editing earlier input, so the cache prefix stays valid.
- Above 272K input tokens the whole request bills input and cache at 2x and output at 1.5x, so long-context costs jump suddenly; compaction or retrieval can keep input under 272K. Through Wokey there is no tier, and the whole 1M context bills at one rate.
- Behavior changes: it asks clarifying questions and pauses for confirmation, delegates less and tests more broadly; formatting is more detailed and phrases repeat. It is more sensitive to skills and context files, so OpenAI recommends auditing existing skills and turning them into minimal routers, since over-specific guidance can hold it back.
- Cybersecurity requests: the public model refuses to write proof-of-concept exploits, so penetration testing and exploit reproduction tasks may be refused outright, while secure code review and patching work normally. Full capability requires Daybreak access.
- In Codex, run $openai-docs migrate this project to GPT-6 Astra to update the project against the official checklist. Fast mode has no latency SLA, and EU data residency supports Standard processing only.
Request body after migration (/v1/responses)
{
"model": "gpt-6-astra",
"reasoning": { "effort": "medium" },
"input": [{ "role": "user", "content": "Find and fix the flaky test in this repo" }]
}- Model ID
gpt-6-astra- Vendors
- openai
- Supply status
- Currently callable
Catalog pricing snapshot
Prices are per 1M tokens; request-time pricing controls settlement.
| Meter | Wokey | Official reference | Savings |
|---|---|---|---|
| Input | $0.9 | $10 | 91% |
| Output | $4.5 | $50 | 91% |
| Cache read | $0.09 | $1 | 91% |
| Cache write | $1.125 | $12.5 | 91% |
Official price source: https://developers.openai.com/api/docs/models/gpt-6-astra
Model capabilities
- Context
- 1,050,000 tokens
- Max output
- 128,000 tokens
- Streaming
- Supported
- Tools
- Supported
- Vision
- Supported
- Supported API forms
/v1/chat/completions·/v1/responses
Send a request
/v1/chat/completions
curl https://api.wokey.ai/v1/chat/completions \
-H "Authorization: Bearer $WOKEY_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"gpt-6-astra","messages":[{"role":"user","content":"Hello"}]}'/v1/responses
curl https://api.wokey.ai/v1/responses \
-H "Authorization: Bearer $WOKEY_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"gpt-6-astra","input":"Hello"}'Use GPT-6 Astra in your tools
The gateway is https://api.wokey.ai (OpenAI-compatible clients usually take https://api.wokey.ai/v1), and the model ID is gpt-6-astra. Open the guide for the client you use.
- Use GPT-6 Astra in Codex CLI: OpenAI Responses · base URL and API key setup
- Use GPT-6 Astra in opencode: OpenAI Chat Completions · base URL and API key setup
- Use GPT-6 Astra in OpenClaw: OpenAI Chat Completions · base URL and API key setup
- Use GPT-6 Astra in Hermes Agent: OpenAI Chat Completions · base URL and API key setup
- OpenRouter alternative: Wokey vs OpenRouter: Compare token prices, payment fees, API forms, and upstream sources.
- Verifiable AI API: check a GPT-6 Astra response came from the official upstream: Which responses carry a tee.proof, and what a proof does and does not show.
- Verify a GPT-6 Astra response online: AI API relay verifier: Paste a response with its tee.proof and check it locally in your browser; nothing is uploaded.
Example production call
These usage values come from a completed GPT-6 Astra call that succeeded on its first attempt and was reconciled with its bill.
- Total input
- 675347
- Output
- 669
- Cached share of input
- 99.1%
- Call duration
- 23.74 s
Responses usage
Compiled from the billing record into API field examples. input_tokens is total input, including cached tokens.
{
"model": "gpt-6-astra",
"usage": {
"input_tokens": 675347,
"input_tokens_details": {
"cached_tokens": 669312
},
"output_tokens": 669,
"total_tokens": 676016
}
}Token accounting
These are the final per-million-token rates recorded for this call.
| Meter | Tokens | Rate for this call / 1M |
|---|---|---|
| Uncached input | 6035 | $0.9 |
| Cache read | 669312 | $0.09 |
| Output | 669 | $4.5 |
- Billed amount
- $0.068680
Cost = sum of tokens × their rate ÷ 1,000,000, rounded to six decimals after summing.
Source of this call
This call used the OpenAI channel at chatgpt.com with OAuth authorization. Usage was metered upstream and the request returned HTTP 200.
About Provider Node · Integration docsVerify Prompt Cache
- Keep the shared prefix of long prompts identical and reuse context according to the selected API’s cache rules.
- Check input_tokens_details.cached_tokens; zero means no cache reads were recorded for this call.
- Uncached input = input_tokens minus cached tokens. Price each bucket separately and add the output charge.
Frequently asked questions
GPT-6 Astra vs Claude Opus 5.5: which is better?
It depends on the task. In Anthropic’s published scores Opus 5.5 leads on Terminal-Bench 4.0 (66.4% vs 57.9%), GDPval-AA (1846 vs 1542) and Humanity’s Last Exam (67.7% vs 57.2%), while GPT-6 Astra leads on AutomationBench (41.4% vs 40.0%) and Terminal-Bench-Science (64.6% vs 58.7%). Opus 5.5 lists at $4 / $20, two-fifths of Astra’s $10 / $50, but Artificial Analysis measured far fewer output tokens for Astra, so a max-effort task costs $3.26 versus $5.98 on Opus 5.5, which scores higher on the index (58 vs 53).
Should I use GPT-6 Astra, GPT-6 Sol or GPT-6 Luna?
OpenAI positions Astra as its most capable model, Sol for demanding tasks and Luna for scale. They list at $10 / $50, $2 / $10 and $0.10 / $0.50. On Artificial Analysis’ Intelligence Index Astra scores 53 at max for $3.26 per task, Sol 48 for $1.06 and Luna 37 for $0.07. Start everyday coding on Sol, move to Astra for problems Sol handles poorly, and use Luna for bulk classification and extraction.
How is GPT-6 Astra billed above 272K context?
At official prices, once input exceeds 272K tokens the whole request bills input and cache at 2x and output at 1.5x: $20 input, $2 cache read, $25 cache write and $75 output per million tokens, versus $10 / $1 / $12.50 / $50 below 272K. Batch and Flex are 50% off and Fast mode is 2x. Through Wokey there is no tier, and the whole 1M context bills at one rate of $0.90 input / $4.50 output per million tokens.
Why does GPT-6 Astra refuse cybersecurity requests, and what is Daybreak?
GPT-6 Astra scored 100% on ExploitBench, and OpenAI rates its cyber capability as Critical under the Preparedness Framework. The public model therefore refuses to write proof-of-concept exploits and is limited to secure code review and patching. Daybreak is OpenAI’s access program for trusted security teams with fuller capability, which OpenAI says will expand. The same limits apply through Wokey, which does not offer Daybreak access.
Does GPT-6 Astra support temperature or reasoning effort none?
No. GPT-6 Astra owns its sampling policy, and OpenAI says to remove temperature, top_p, top_logprobs and logprobs; reasoning effort accepts only low, medium, high, xhigh and max, so replace none with low. Through Wokey the gateway strips these sampling fields and maps none and minimal to low, so older requests work unchanged.
GPT-6 Astra uses up my Codex subscription too fast. What can I do?
GPT-6 Astra usage in ChatGPT and Codex counts toward subscription allowances, and Hacker News users report it burning about 2.5x the rate of GPT-5.6 Sol, with one using a 5-hour limit in 15 messages. Lower effort to medium, hand routine work to GPT-6 Sol, or switch to per-token API billing. Through Wokey it costs $0.90 input / $4.50 output per million tokens (official $10 / $50); set base_url to https://api.wokey.ai, wire_api to responses and model to gpt-6-astra in Codex’s config.toml.
How much does the GPT-6 Astra API cost, and how much does it save versus the official API?
Wokey lists input at $0.9 per 1M tokens and output at $4.5 per 1M tokens; the official references are $10 and $50. That is 91% lower for input and 91% lower for output. Request-time pricing controls settlement, and the vendor pricing source is linked on this page.
What is the GPT-6 Astra API model ID?
The model ID for GPT-6 Astra on Wokey is gpt-6-astra. Use it exactly as written in the request body's model field or in your client's model setting (for example ANTHROPIC_MODEL in Claude Code, or model in the Codex CLI config.toml); GET https://api.wokey.ai/v1/models also lists it.
Which API forms does GPT-6 Astra support, and how do I call it through Wokey?
GPT-6 Astra supports /v1/chat/completions (Chat Completions requests) and /v1/responses (Responses requests). Set the API base URL to https://api.wokey.ai, authenticate with a Wokey API key, and use the canonical model ID gpt-6-astra; you can start with /v1/responses. The Send a request section includes a curl example for every supported API form.
How do I use GPT-6 Astra in Codex CLI?
Codex CLI calls GPT-6 Astra over the Responses protocol. Add a model_provider to ~/.codex/config.toml with base_url = "https://api.wokey.ai", wire_api = "responses", and env_key = "OPENAI_API_KEY"; set model to gpt-6-astra and put your Wokey API key in OPENAI_API_KEY. The Codex CLI guide has the full config.
What are the context window and maximum output for GPT-6 Astra?
The catalog context window is 1,050,000 tokens and the maximum output is 128,000 tokens. These are separate limits, and a request must also satisfy the upstream constraints in effect when it is processed.
How are cache reads and cache writes priced for GPT-6 Astra?
Cache read: Wokey lists $0.09 per 1M tokens and the official reference is $1 per 1M tokens. Cache write: Wokey lists $1.125 per 1M tokens and the official reference is $12.5 per 1M tokens.
How does calling GPT-6 Astra through Wokey differ from OpenRouter?
Wokey lists GPT-6 Astra at $0.9 input / $4.5 output per 1M tokens, against an official reference of $10 / $50. OpenRouter states it adds no markup on inference and charges its fee when you buy credits (5.5% by card). Wokey carries a curated set of models and has no per-provider routing parameters, so OpenRouter fits better if you need a wider catalog. The OpenRouter alternative page has the full comparison.
How do GPT-6 Astra and GPT-6.1 Sol differ in price and context?
In the published price comparison, GPT-6 Astra lists input/output at $0.9 / $4.5 with a 1,050,000-token context window; GPT-6.1 Sol lists $0.4 / $2 with a 1,050,000-token context window. This is a factual price-and-capacity comparison, not a quality ranking; also compare their supported API forms and capabilities.
Related models and documentation
Get API key · gpt-6-astra