gpt-oss-20b

OpenAI · Active · available in 13 of 31 regions

Batch inferenceStreaming Tool useVisionPrompt cachingEmbeddingsFine-tuning
Context window
Max output
Input / 1M
$0.070
us-east-1
Output / 1M
$0.300

What gpt-oss-20b is good for

editorial

gpt-oss-20b, and is the cheapest option from this provider at $0.070 per 1M input tokens. Callable in 13 of 31 regions.

Suits

  • Bulk jobs. Batch inference is available, typically at half the on-demand rate, for work that can wait — backfills, nightly enrichment, evaluation runs.
  • Cost. The cheapest OpenAI model here at $0.070 per 1M input tokens — the default choice for high-volume, low-difficulty work.

Think twice if

  • Agentic loops. No prompt caching published, so every step re-pays full price for the prompt prefix. Expensive for agents that resend a long context.
  • Function calling. No tool use, so it cannot drive an agent loop or call your APIs — it answers from what is in the prompt.
  • Documents with layout. Text only. Scanned PDFs, screenshots and charts need a vision model or an OCR step first.

This section is judgement, not data from AWS. It is composed from the capabilities, context window, price position and region coverage shown elsewhere on this page — so it stays in step with the daily snapshot rather than going stale.

Where teams typically use it

Chat and assistants

Conversational work where responsiveness matters as much as depth.

  • · A customer-facing assistant where the first token needs to appear in under a second.
  • · An in-product copilot that explains what the user is looking at and answers follow-ups.
  • · A triage bot that qualifies an incoming request before routing it to the right team.

Coding

Writing, reviewing and refactoring code, usually across more than one file.

  • · A review pass that flags real defects with a severity and leaves style alone.
  • · A framework upgrade applied across a repository, one module at a time, tests green between each.
  • · Turning a failing bug report into a reproducing test and then a fix.

Model & inference profile IDs

Base model ID — direct on-demand invoke

Pricing by region

USD per 1M tokens
Region Input Output Cache read Cache write Batch in Batch out Source
us-east-1 US East (N. Virginia) $0.070 $0.300 $0.035 $0.150 API
us-east-2 US East (Ohio) $0.070 $0.300 $0.035 $0.150 API
us-west-2 US West (Oregon) $0.070 $0.300 $0.035 $0.150 API
eu-west-1 EU (Ireland) $0.080 $0.350 $0.040 $0.180 API
eu-west-2 EU (London) $0.110 $0.470 $0.050 $0.230 API
eu-central-1 EU (Frankfurt) $0.090 $0.400 $0.050 $0.200 API
eu-north-1 EU (Stockholm) $0.070 $0.300 $0.035 $0.150 API
eu-south-1 EU (Milan) $0.090 $0.400 $0.050 $0.200 API
ap-northeast-1 Asia Pacific (Tokyo) $0.080 $0.360 $0.040 $0.180 API
ap-south-1 Asia Pacific (Mumbai) $0.080 $0.350 $0.040 $0.180 API
ap-southeast-2 Asia Pacific (Sydney) $0.072 $0.309 $0.036 $0.154 API
ap-southeast-3 Asia Pacific (Jakarta) $0.070 $0.310 $0.040 $0.160 API
sa-east-1 South America (Sao Paulo) $0.080 $0.360 $0.040 $0.180 API

Region availability

TPM / RPM are default account quotas from AWS Service Quotas where a model-specific limit is published. They are per-account defaults and adjustable on request.

Where this model actually runs

Under a strict residency constraint, gpt-oss-20b can be served without the request leaving 12 of the 25 jurisdictions with a Bedrock region.

gpt-oss-20b appears in

Common questions

How much does gpt-oss-20b cost on Amazon Bedrock?
$0.070 per 1M input tokens and $0.300 per 1M output tokens in us-east-1.
Which AWS regions support gpt-oss-20b?
13 of 31 regions: us-east-1, us-east-2, us-west-2, eu-west-1, eu-west-2, eu-central-1, eu-north-1, eu-south-1, ap-northeast-1, ap-south-1, ap-southeast-2, ap-southeast-3, sa-east-1.
What is the model ID for gpt-oss-20b on Bedrock?
openai.gpt-oss-20b-1:0.

Other OpenAI models