GPT-5.6 Luna

OpenAI · Active · available in 28 of 31 regions

Tool useVisionPrompt cachingStreaming Batch inferenceEmbeddingsFine-tuning
Context window
1M
1,000,000 tokens
Max output
Input / 1M
$0.200
us-east-1
Output / 1M
$1.20

What GPT-5.6 Luna is good for

editorial

GPT-5.6 Luna offers a 1M-token context window with tool use and prompt caching at $0.200 per 1M input tokens. Callable in 28 of 31 regions.

Suits

  • Very long inputs. A 1M-token window holds an entire codebase or a day of transcripts in one call, so you can skip chunking.
  • Repeated prompts. Supports prompt caching, so a re-sent prefix costs about 10× less than fresh input. On agentic loops that resend a long system prompt every step, this moves the bill more than the headline price does.
  • Tool use. Can call functions, so it can drive retrieval, look things up, and take actions rather than only answering.
  • Images. Accepts image input: screenshots, scanned documents, charts and UI captures.

Think twice if

  • Invocation. Reachable only through a cross-region inference profile — calling the bare model ID returns a validation error. Use an ID such as global.openai.gpt-5.6-luna.

This section is judgement, not data from AWS. It is composed from the capabilities, context window, price position and region coverage shown elsewhere on this page — so it stays in step with the daily snapshot rather than going stale.

Where teams typically use it

Chat and assistants

Conversational work where responsiveness matters as much as depth.

  • · A customer-facing assistant where the first token needs to appear in under a second.
  • · An in-product copilot that explains what the user is looking at and answers follow-ups.
  • · A triage bot that qualifies an incoming request before routing it to the right team.

Summarisation and extraction

High-volume, structured output from unstructured input. Usually the cheapest workload to run well.

  • · Turning a million support tickets into a fixed schema of product, severity and root cause.
  • · Pulling line items, totals and dates off scanned invoices into JSON.
  • · Condensing every meeting transcript into decisions taken and owners assigned.

Model & inference profile IDs

Base model ID — direct on-demand invoke

Cross-region inference profiles (3)

Regions marked Inference profile only require a profile ID rather than the base model ID — invoking the base ID there returns a validation error. Regional (us., eu.) profiles carry a 10% premium over global.

Pricing by region

USD per 1M tokens
Region Input Output Cache read Cache write Batch in Batch out Source
us-east-1 US East (N. Virginia) $0.200 $1.20 $0.020 $0.250 curated
us-east-2 US East (Ohio) $0.200 $1.20 $0.020 $0.250 curated
us-west-1 US West (N. California) $0.200 $1.20 $0.020 $0.250 curated
us-west-2 US West (Oregon) $0.200 $1.20 $0.020 $0.250 curated
ca-central-1 Canada (Central) $0.200 $1.20 $0.020 $0.250 curated
ca-west-1 Canada West (Calgary) $0.200 $1.20 $0.020 $0.250 curated
eu-west-1 EU (Ireland) $0.200 $1.20 $0.020 $0.250 curated
eu-west-2 EU (London) $0.200 $1.20 $0.020 $0.250 curated
eu-west-3 EU (Paris) $0.200 $1.20 $0.020 $0.250 curated
eu-central-1 EU (Frankfurt) $0.200 $1.20 $0.020 $0.250 curated
eu-central-2 Europe (Zurich) $0.200 $1.20 $0.020 $0.250 curated
eu-north-1 EU (Stockholm) $0.200 $1.20 $0.020 $0.250 curated
eu-south-1 EU (Milan) $0.200 $1.20 $0.020 $0.250 curated
eu-south-2 Europe (Spain) $0.200 $1.20 $0.020 $0.250 curated
ap-east-2 Asia Pacific (Taipei) $0.200 $1.20 $0.020 $0.250 curated
ap-northeast-1 Asia Pacific (Tokyo) $0.200 $1.20 $0.020 $0.250 curated
ap-northeast-2 Asia Pacific (Seoul) $0.200 $1.20 $0.020 $0.250 curated
ap-northeast-3 Asia Pacific (Osaka) $0.200 $1.20 $0.020 $0.250 curated
ap-south-1 Asia Pacific (Mumbai) $0.200 $1.20 $0.020 $0.250 curated
ap-south-2 Asia Pacific (Hyderabad) $0.200 $1.20 $0.020 $0.250 curated
ap-southeast-1 Asia Pacific (Singapore) $0.200 $1.20 $0.020 $0.250 curated
ap-southeast-2 Asia Pacific (Sydney) $0.200 $1.20 $0.020 $0.250 curated
ap-southeast-3 Asia Pacific (Jakarta) $0.200 $1.20 $0.020 $0.250 curated
ap-southeast-4 Asia Pacific (Melbourne) $0.200 $1.20 $0.020 $0.250 curated
ap-southeast-5 Asia Pacific (Malaysia) $0.200 $1.20 $0.020 $0.250 curated
ap-southeast-7 Asia Pacific (Thailand) $0.200 $1.20 $0.020 $0.250 curated
il-central-1 Israel (Tel Aviv) $0.200 $1.20 $0.020 $0.250 curated
sa-east-1 South America (Sao Paulo) $0.200 $1.20 $0.020 $0.250 curated

Curated pricing. The AWS Price List API publishes no SKU for this model, so these figures are hand-maintained from AWS's public pricing page and verified on Aug 27, 2026. They are not machine-derived — confirm in the AWS console before committing spend. Global endpoint, requests up to 272K tokens. Beyond that AWS bills 2x input and 1.5x output.

Region availability

TPM / RPM are default account quotas from AWS Service Quotas where a model-specific limit is published. They are per-account defaults and adjustable on request.

Where this model actually runs

Under a strict residency constraint, GPT-5.6 Luna can be served without the request leaving 2 of the 25 jurisdictions with a Bedrock region.

GPT-5.6 Luna compared

Side-by-side pages against every model people weigh this one against.

GPT-5.6 Luna vs Claude Fable 5GPT-5.6 Luna costs 50× less per input token than Claude Fable 5.GPT-5.6 Luna vs Claude Haiku 4.5GPT-5.6 Luna costs 5.0× less per input token than Claude Haiku 4.5, and GPT-5.6 Luna holds 1M tokens against 200K for Claude Haiku 4.5.GPT-5.6 Luna vs Claude Opus 4.5GPT-5.6 Luna costs 25× less per input token than Claude Opus 4.5, and GPT-5.6 Luna holds 1M tokens against 200K for Claude Opus 4.5.GPT-5.6 Luna vs Claude Opus 4.6GPT-5.6 Luna costs 25× less per input token than Claude Opus 4.6.GPT-5.6 Luna vs Claude Opus 4.7GPT-5.6 Luna costs 25× less per input token than Claude Opus 4.7.GPT-5.6 Luna vs Claude Opus 4.8GPT-5.6 Luna costs 25× less per input token than Claude Opus 4.8.GPT-5.6 Luna vs Claude Opus 5GPT-5.6 Luna costs 25× less per input token than Claude Opus 5.GPT-5.6 Luna vs Claude Sonnet 4.5GPT-5.6 Luna costs 15× less per input token than Claude Sonnet 4.5, and GPT-5.6 Luna holds 1M tokens against 200K for Claude Sonnet 4.5.GPT-5.6 Luna vs Claude Sonnet 4.6GPT-5.6 Luna costs 15× less per input token than Claude Sonnet 4.6.GPT-5.6 Luna vs Claude Sonnet 5GPT-5.6 Luna costs 15× less per input token than Claude Sonnet 5.GPT-5.6 Luna vs GPT-5.6 SolGPT-5.6 Luna costs 20× less per input token than GPT-5.6 Sol.GPT-5.6 Luna vs GPT-5.6 TerraGPT-5.6 Luna costs 10× less per input token than GPT-5.6 Terra.GPT-5.6 Luna vs Grok 4.6GPT-5.6 Luna costs 10× less per input token than Grok 4.6, and GPT-5.6 Luna holds 1M tokens against 500K for Grok 4.6.GPT-5.6 Luna vs Nova 2 LiteGPT-5.6 Luna costs 1.5× less per input token than Nova 2 Lite, and only GPT-5.6 Luna does tool use (Nova 2 Lite does not).GPT-5.6 Luna vs Nova LiteNova Lite costs 3.3× less per input token than GPT-5.6 Luna, and GPT-5.6 Luna holds 1M tokens against 300K for Nova Lite.

GPT-5.6 Luna appears in

Common questions

How much does GPT-5.6 Luna cost on Amazon Bedrock?
$0.200 per 1M input tokens and $1.20 per 1M output tokens in us-east-1. Cached input reads at $0.020, roughly a tenth of the input rate, which dominates the bill on agentic workloads that re-send a long prefix.
Which AWS regions support GPT-5.6 Luna?
28 of 31 regions: us-east-1, us-east-2, us-west-1, us-west-2, ca-central-1, ca-west-1, eu-west-1, eu-west-2, eu-west-3, eu-central-1, eu-central-2, eu-north-1, eu-south-1, eu-south-2, ap-east-2, ap-northeast-1, ap-northeast-2, ap-northeast-3, ap-south-1, ap-south-2, ap-southeast-1, ap-southeast-2, ap-southeast-3, ap-southeast-4, ap-southeast-5, ap-southeast-7, il-central-1, sa-east-1. Regions marked profile-only need a cross-region inference profile ID such as global.openai.gpt-5.6-luna rather than the bare model ID.
What is the context window of GPT-5.6 Luna?
1,000,000 tokens (1M).
What is the model ID for GPT-5.6 Luna on Bedrock?
openai.gpt-5.6-luna. In regions where it is only reachable through a cross-region inference profile, use a prefixed ID instead: global.openai.gpt-5.6-luna, in.openai.gpt-5.6-luna, us.openai.gpt-5.6-luna.

Other OpenAI models