GPT-5.6 Luna
OpenAI · Active · available in 28 of 31 regions
- Context window
- 1M
- 1,000,000 tokens
- Max output
- —
- Input / 1M
- $0.200
- us-east-1
- Output / 1M
- $1.20
What GPT-5.6 Luna is good for
editorialGPT-5.6 Luna offers a 1M-token context window with tool use and prompt caching at $0.200 per 1M input tokens. Callable in 28 of 31 regions.
Suits
- Very long inputs. A 1M-token window holds an entire codebase or a day of transcripts in one call, so you can skip chunking.
- Repeated prompts. Supports prompt caching, so a re-sent prefix costs about 10× less than fresh input. On agentic loops that resend a long system prompt every step, this moves the bill more than the headline price does.
- Tool use. Can call functions, so it can drive retrieval, look things up, and take actions rather than only answering.
- Images. Accepts image input: screenshots, scanned documents, charts and UI captures.
Think twice if
- Invocation. Reachable only through a cross-region inference profile — calling the bare model ID returns a validation error. Use an ID such as global.openai.gpt-5.6-luna.
This section is judgement, not data from AWS. It is composed from the capabilities, context window, price position and region coverage shown elsewhere on this page — so it stays in step with the daily snapshot rather than going stale.
Where teams typically use it
Chat and assistants
Conversational work where responsiveness matters as much as depth.
- · A customer-facing assistant where the first token needs to appear in under a second.
- · An in-product copilot that explains what the user is looking at and answers follow-ups.
- · A triage bot that qualifies an incoming request before routing it to the right team.
Summarisation and extraction
High-volume, structured output from unstructured input. Usually the cheapest workload to run well.
- · Turning a million support tickets into a fixed schema of product, severity and root cause.
- · Pulling line items, totals and dates off scanned invoices into JSON.
- · Condensing every meeting transcript into decisions taken and owners assigned.
Model & inference profile IDs
Base model ID — direct on-demand invoke
Cross-region inference profiles (3)
Regions marked Inference profile only require a profile ID rather than the base model ID — invoking the base ID there returns a validation error. Regional (us., eu.) profiles carry a 10% premium over global.
Pricing by region
USD per 1M tokens| Region | Input | Output | Cache read | Cache write | Batch in | Batch out | Source |
|---|---|---|---|---|---|---|---|
| us-east-1 US East (N. Virginia) | $0.200 | $1.20 | $0.020 | $0.250 | — | — | curated |
| us-east-2 US East (Ohio) | $0.200 | $1.20 | $0.020 | $0.250 | — | — | curated |
| us-west-1 US West (N. California) | $0.200 | $1.20 | $0.020 | $0.250 | — | — | curated |
| us-west-2 US West (Oregon) | $0.200 | $1.20 | $0.020 | $0.250 | — | — | curated |
| ca-central-1 Canada (Central) | $0.200 | $1.20 | $0.020 | $0.250 | — | — | curated |
| ca-west-1 Canada West (Calgary) | $0.200 | $1.20 | $0.020 | $0.250 | — | — | curated |
| eu-west-1 EU (Ireland) | $0.200 | $1.20 | $0.020 | $0.250 | — | — | curated |
| eu-west-2 EU (London) | $0.200 | $1.20 | $0.020 | $0.250 | — | — | curated |
| eu-west-3 EU (Paris) | $0.200 | $1.20 | $0.020 | $0.250 | — | — | curated |
| eu-central-1 EU (Frankfurt) | $0.200 | $1.20 | $0.020 | $0.250 | — | — | curated |
| eu-central-2 Europe (Zurich) | $0.200 | $1.20 | $0.020 | $0.250 | — | — | curated |
| eu-north-1 EU (Stockholm) | $0.200 | $1.20 | $0.020 | $0.250 | — | — | curated |
| eu-south-1 EU (Milan) | $0.200 | $1.20 | $0.020 | $0.250 | — | — | curated |
| eu-south-2 Europe (Spain) | $0.200 | $1.20 | $0.020 | $0.250 | — | — | curated |
| ap-east-2 Asia Pacific (Taipei) | $0.200 | $1.20 | $0.020 | $0.250 | — | — | curated |
| ap-northeast-1 Asia Pacific (Tokyo) | $0.200 | $1.20 | $0.020 | $0.250 | — | — | curated |
| ap-northeast-2 Asia Pacific (Seoul) | $0.200 | $1.20 | $0.020 | $0.250 | — | — | curated |
| ap-northeast-3 Asia Pacific (Osaka) | $0.200 | $1.20 | $0.020 | $0.250 | — | — | curated |
| ap-south-1 Asia Pacific (Mumbai) | $0.200 | $1.20 | $0.020 | $0.250 | — | — | curated |
| ap-south-2 Asia Pacific (Hyderabad) | $0.200 | $1.20 | $0.020 | $0.250 | — | — | curated |
| ap-southeast-1 Asia Pacific (Singapore) | $0.200 | $1.20 | $0.020 | $0.250 | — | — | curated |
| ap-southeast-2 Asia Pacific (Sydney) | $0.200 | $1.20 | $0.020 | $0.250 | — | — | curated |
| ap-southeast-3 Asia Pacific (Jakarta) | $0.200 | $1.20 | $0.020 | $0.250 | — | — | curated |
| ap-southeast-4 Asia Pacific (Melbourne) | $0.200 | $1.20 | $0.020 | $0.250 | — | — | curated |
| ap-southeast-5 Asia Pacific (Malaysia) | $0.200 | $1.20 | $0.020 | $0.250 | — | — | curated |
| ap-southeast-7 Asia Pacific (Thailand) | $0.200 | $1.20 | $0.020 | $0.250 | — | — | curated |
| il-central-1 Israel (Tel Aviv) | $0.200 | $1.20 | $0.020 | $0.250 | — | — | curated |
| sa-east-1 South America (Sao Paulo) | $0.200 | $1.20 | $0.020 | $0.250 | — | — | curated |
Curated pricing. The AWS Price List API publishes no SKU for this model, so these figures are hand-maintained from AWS's public pricing page and verified on Aug 27, 2026. They are not machine-derived — confirm in the AWS console before committing spend. Global endpoint, requests up to 272K tokens. Beyond that AWS bills 2x input and 1.5x output.
Region availability
TPM / RPM are default account quotas from AWS Service Quotas where a model-specific limit is published. They are per-account defaults and adjustable on request.
Where this model actually runs
Under a strict residency constraint, GPT-5.6 Luna can be served without the request leaving 2 of the 25 jurisdictions with a Bedrock region.
GPT-5.6 Luna compared
Side-by-side pages against every model people weigh this one against.
GPT-5.6 Luna appears in
Common questions
- How much does GPT-5.6 Luna cost on Amazon Bedrock?
- $0.200 per 1M input tokens and $1.20 per 1M output tokens in us-east-1. Cached input reads at $0.020, roughly a tenth of the input rate, which dominates the bill on agentic workloads that re-send a long prefix.
- Which AWS regions support GPT-5.6 Luna?
- 28 of 31 regions: us-east-1, us-east-2, us-west-1, us-west-2, ca-central-1, ca-west-1, eu-west-1, eu-west-2, eu-west-3, eu-central-1, eu-central-2, eu-north-1, eu-south-1, eu-south-2, ap-east-2, ap-northeast-1, ap-northeast-2, ap-northeast-3, ap-south-1, ap-south-2, ap-southeast-1, ap-southeast-2, ap-southeast-3, ap-southeast-4, ap-southeast-5, ap-southeast-7, il-central-1, sa-east-1. Regions marked profile-only need a cross-region inference profile ID such as global.openai.gpt-5.6-luna rather than the bare model ID.
- What is the context window of GPT-5.6 Luna?
- 1,000,000 tokens (1M).
- What is the model ID for GPT-5.6 Luna on Bedrock?
- openai.gpt-5.6-luna. In regions where it is only reachable through a cross-region inference profile, use a prefixed ID instead: global.openai.gpt-5.6-luna, in.openai.gpt-5.6-luna, us.openai.gpt-5.6-luna.