Estimate GPT-6 Astra, Sol and Luna alongside GPT-5.6 API costs, including long-context pricing, cache writes, cache reads and output tokens. Enter representative per-request token counts and choose Standard, Batch/Flex or Fast processing. The calculator applies OpenAI’s higher long-context matrix automatically above 272,000 input tokens.
Use token counts from a representative API response wherever possible. Model and processing-mode availability may differ between accounts and regions. Pricing can change; confirm material budgets against OpenAI’s official API pricing and GPT-6 Astra model page, plus the GPT-6 Sol and GPT-6 Luna model pages.
Estimate your OpenAI API cost
Pricing verified: 23 September 2026. Published API rates in USD.
GPT-6 Astra · Standard · Short context · 1,000 requests
- Uncached input
- $10.00
- Cache writes
- $0.00
- Cache reads
- $0.00
- Output
- $25.00
- Average per request
- $0.035
Planning estimate only. It excludes tools, storage, fine-tuning, regional charges, retries, taxes, contracted discounts and non-token fees.
Current OpenAI API rates used by this calculator
Rates below are USD per million tokens for Standard API requests. Batch and Flex processing cost 50% of Standard; Fast processing costs twice the applicable Standard rate. A request with more than 272,000 total input tokens uses the long-context rates for the full request.
| Model and context | Input | Cache read | Cache write | Output |
|---|---|---|---|---|
| GPT-6 Astra – short | $10 | $1 | $12.50 | $50 |
| GPT-6 Astra – long | $20 | $2 | $25 | $75 |
| GPT-6 Sol: short | $2 | $0.20 | $2.50 | $10 |
| GPT-6 Sol: long | $4 | $0.40 | $5 | $15 |
| GPT-6 Luna: short | $0.10 | $0.01 | $0.125 | $0.50 |
| GPT-6 Luna: long | $0.20 | $0.02 | $0.25 | $0.75 |
| GPT-5.6 Sol – short | $4 | $0.40 | $5 | $20 |
| GPT-5.6 Sol – long | $8 | $0.80 | $10 | $30 |
| GPT-5.6 Terra – short | $2 | $0.20 | $2.50 | $12 |
| GPT-5.6 Terra – long | $4 | $0.40 | $5 | $18 |
| GPT-5.6 Luna – short | $0.20 | $0.02 | $0.25 | $1.20 |
| GPT-5.6 Luna – long | $0.40 | $0.04 | $0.50 | $1.80 |
“Long” means more than 272,000 input tokens in one request, not 272,000 tokens across all monthly requests. Cache-write rates equal 1.25 times the corresponding uncached-input rate. GPT-5.6 Sol’s listed rates are OpenAI’s promotional rates, available through at least 21 November 2026.
How to estimate an OpenAI API bill
- Choose the exact model. Check the full generation and model name. GPT-6 Sol and Luna have different prices from GPT-5.6 Sol and Luna.
- Choose the processing mode. Standard is the baseline. Batch/Flex halves the applicable rate, while Fast doubles it.
- Enter per-request token counts. Keep uncached input, cache creation, cache reads and output separate.
- Enter request volume. For a monthly forecast, use the number of calls expected in a typical month.
- Check the context notice. The calculator totals all input categories per request and switches the entire request to the published long-context matrix above 272,000 tokens.
- Reconcile with actual usage. Re-run the estimate when prompts, caching or models change.
Formula used
Total = requests × processing multiplier × ((uncached input × input rate) + (cache writes × write rate) + (cache reads × cached rate) + (output × output rate)) / 1,000,000
Which OpenAI model should you price?
- GPT-6 Astra is the higher-priced option for demanding reasoning and agent workflows.
- GPT-6 Sol supports coding and agent workflows at $2 per million short-context input tokens and $10 per million output tokens.
- GPT-6 Luna supports focused, high-volume tasks at $0.10 per million short-context input tokens and $0.50 per million output tokens.
- GPT-5.6 Sol is an advanced option for difficult workloads and currently has promotional API pricing.
- GPT-5.6 Terra is the balanced option for production agents and general-purpose work.
- GPT-5.6 Luna is the cost-sensitive option for high-volume or simpler workloads.
Model choice should be based on your own quality, latency and reliability tests. A cheaper rate does not guarantee a lower end-to-end cost if a workflow needs more retries or tokens.
Frequently asked questions
Are GPT-6 Astra, Sol and Luna available through the API?
OpenAI lists gpt-6-astra, gpt-6-sol and gpt-6-luna as API models. Sol and Luna were released on 22 September 2026. Check model access and processing-mode availability in your own account before budgeting a deployment.
Are ChatGPT subscriptions included?
No. ChatGPT plans and API usage are billed separately. This tool estimates API token charges only.
Why are cache writes separate from cached input?
Creating a reusable cache and reading from that cache have different rates. Combining them can materially understate a workload that frequently creates new cache entries.
What is the difference between Standard, Batch/Flex and Fast?
Standard uses the published base rates. Batch and Flex processing are priced at 50% of Standard, while Fast processing is twice the applicable Standard rate. Availability and latency characteristics differ, so choose the mode your workload actually uses. GPT-6 Astra Fast is unavailable with EU data residency. GPT-6 Sol and Luna support EU data residency only with Standard processing. The 10% regional-processing surcharge is excluded from this calculator.
Why might the final bill differ?
Retries, tool calls, hidden system or tool-schema tokens, model changes, storage, regional pricing, contract terms and taxes can all change the final invoice.
Compare providers with the Claude API cost calculator and Gemini API cost calculator, or visit the AI cost calculators hub.