Claude 3 Sonnet

Anthropic · Active · available in 13 of 31 regions

Tool useVisionStreaming Prompt cachingBatch inferenceEmbeddingsFine-tuning
Context window
200K
200,000 tokens
Max output
Input / 1M
$3.00
us-east-1
Output / 1M

What Claude 3 Sonnet is good for

editorial

Claude 3 Sonnet offers a 200K-token context window with tool use at $3.00 per 1M input tokens. Callable in 13 of 31 regions.

Suits

  • Long documents. 200K tokens covers long documents and multi-file context without splitting them.
  • Tool use. Can call functions, so it can drive retrieval, look things up, and take actions rather than only answering.
  • Images. Accepts image input: screenshots, scanned documents, charts and UI captures.

Think twice if

  • Agentic loops. No prompt caching published, so every step re-pays full price for the prompt prefix. Expensive for agents that resend a long context.

This section is judgement, not data from AWS. It is composed from the capabilities, context window, price position and region coverage shown elsewhere on this page — so it stays in step with the daily snapshot rather than going stale.

Where teams typically use it

Chat and assistants

Conversational work where responsiveness matters as much as depth.

  • · A customer-facing assistant where the first token needs to appear in under a second.
  • · An in-product copilot that explains what the user is looking at and answers follow-ups.
  • · A triage bot that qualifies an incoming request before routing it to the right team.

Model & inference profile IDs

Base model ID — direct on-demand invoke

Cross-region inference profiles (3)

Regions marked Inference profile only require a profile ID rather than the base model ID — invoking the base ID there returns a validation error. Regional (us., eu.) profiles carry a 10% premium over global.

Pricing by region

USD per 1M tokens
Region Input Output Cache read Cache write Batch in Batch out Source
us-east-1 US East (N. Virginia) $3.00 API
us-west-2 US West (Oregon) $3.00 API

Region availability

TPM / RPM are default account quotas from AWS Service Quotas where a model-specific limit is published. They are per-account defaults and adjustable on request.

Where this model actually runs

Under a strict residency constraint, Claude 3 Sonnet can be served without the request leaving 6 of the 25 jurisdictions with a Bedrock region.

Claude 3 Sonnet appears in

Common questions

How much does Claude 3 Sonnet cost on Amazon Bedrock?
$3.00 per 1M input tokens and — per 1M output tokens in us-east-1.
Which AWS regions support Claude 3 Sonnet?
13 of 31 regions: us-east-1, us-west-2, ca-central-1, eu-west-1, eu-west-2, eu-west-3, eu-central-1, ap-northeast-1, ap-northeast-2, ap-south-1, ap-southeast-1, ap-southeast-2, sa-east-1. Regions marked profile-only need a cross-region inference profile ID such as apac.anthropic.claude-3-sonnet-20240229-v1:0 rather than the bare model ID.
What is the context window of Claude 3 Sonnet?
200,000 tokens (200K).
What is the model ID for Claude 3 Sonnet on Bedrock?
anthropic.claude-3-sonnet-20240229-v1:0. In regions where it is only reachable through a cross-region inference profile, use a prefixed ID instead: apac.anthropic.claude-3-sonnet-20240229-v1:0, eu.anthropic.claude-3-sonnet-20240229-v1:0, us.anthropic.claude-3-sonnet-20240229-v1:0.

Other Anthropic models