Claude Sonnet 4

Anthropic · LEGACY · available in 23 of 31 regions

Tool useVisionPrompt cachingStreaming Batch inferenceEmbeddingsFine-tuning
Context window
200K
200,000 tokens
Max output
Input / 1M
Output / 1M

What Claude Sonnet 4 is good for

editorial

Claude Sonnet 4 offers a 200K-token context window with tool use and prompt caching, with no token price published by AWS. Callable in 23 of 31 regions.

Suits

  • Long documents. 200K tokens covers long documents and multi-file context without splitting them.
  • Repeated prompts. Supports prompt caching. On agentic loops that resend a long system prompt every step, this moves the bill more than the headline price does.
  • Tool use. Can call functions, so it can drive retrieval, look things up, and take actions rather than only answering.
  • Images. Accepts image input: screenshots, scanned documents, charts and UI captures.

Think twice if

  • Invocation. Reachable only through a cross-region inference profile — calling the bare model ID returns a validation error. Use an ID such as apac.anthropic.claude-sonnet-4-20250514-v1:0.

This section is judgement, not data from AWS. It is composed from the capabilities, context window, price position and region coverage shown elsewhere on this page — so it stays in step with the daily snapshot rather than going stale.

Where teams typically use it

Chat and assistants

Conversational work where responsiveness matters as much as depth.

  • · A customer-facing assistant where the first token needs to appear in under a second.
  • · An in-product copilot that explains what the user is looking at and answers follow-ups.
  • · A triage bot that qualifies an incoming request before routing it to the right team.

Coding

Writing, reviewing and refactoring code, usually across more than one file.

  • · A review pass that flags real defects with a severity and leaves style alone.
  • · A framework upgrade applied across a repository, one module at a time, tests green between each.
  • · Turning a failing bug report into a reproducing test and then a fix.

Model & inference profile IDs

Base model ID — direct on-demand invoke

Cross-region inference profiles (4)

Regions marked Inference profile only require a profile ID rather than the base model ID — invoking the base ID there returns a validation error. Regional (us., eu.) profiles carry a 10% premium over global.

Pricing by region

USD per 1M tokens

No published token pricing

The AWS Price List API publishes no token SKU for this model, and it is not in our curated overlay. This is common for embedding, image and video models, which bill per request or per second rather than per token.

Check the AWS pricing page →

Region availability

TPM / RPM are default account quotas from AWS Service Quotas where a model-specific limit is published. They are per-account defaults and adjustable on request.

Where this model actually runs

Under a strict residency constraint, Claude Sonnet 4 can be served without the request leaving 2 of the 25 jurisdictions with a Bedrock region.

Claude Sonnet 4 appears in

Common questions

Which AWS regions support Claude Sonnet 4?
23 of 31 regions: us-east-1, us-east-2, us-west-1, us-west-2, eu-west-1, eu-west-3, eu-central-1, eu-north-1, eu-south-1, eu-south-2, ap-east-2, ap-northeast-1, ap-northeast-2, ap-northeast-3, ap-south-1, ap-south-2, ap-southeast-1, ap-southeast-2, ap-southeast-3, ap-southeast-4, ap-southeast-5, ap-southeast-7, il-central-1. Regions marked profile-only need a cross-region inference profile ID such as apac.anthropic.claude-sonnet-4-20250514-v1:0 rather than the bare model ID.
What is the context window of Claude Sonnet 4?
200,000 tokens (200K).
What is the model ID for Claude Sonnet 4 on Bedrock?
anthropic.claude-sonnet-4-20250514-v1:0. In regions where it is only reachable through a cross-region inference profile, use a prefixed ID instead: apac.anthropic.claude-sonnet-4-20250514-v1:0, eu.anthropic.claude-sonnet-4-20250514-v1:0, global.anthropic.claude-sonnet-4-20250514-v1:0.

Other Anthropic models