Llama 3.1 70B Instruct

Meta · Active · available in 3 of 16 regions

Tool useBatch inferenceStreaming VisionPrompt cachingEmbeddingsFine-tuning
Context window
128K
128,000 tokens
Max output
Input / 1M
$0.720
us-east-1
Output / 1M
$0.720

Model & inference profile IDs

Base model ID — direct on-demand invoke

Cross-region inference profiles (1)

Regions marked Inference profile only require a profile ID rather than the base model ID — invoking the base ID there returns a validation error. Regional (us., eu.) profiles carry a 10% premium over global.

Pricing by region

USD per 1M tokens
Region Input Output Cache read Cache write Batch in Batch out Source
us-east-1 US East (N. Virginia) $0.720 $0.720 $0.360 $0.360 API
us-east-2 US East (Ohio) $0.720 $0.720 $0.360 $0.360 API
us-west-2 US West (Oregon) $0.720 $0.720 $0.360 $0.360 API

Region availability

TPM / RPM are default account quotas from AWS Service Quotas where a model-specific limit is published. They are per-account defaults and adjustable on request.

Common questions

How much does Llama 3.1 70B Instruct cost on Amazon Bedrock?
$0.720 per 1M input tokens and $0.720 per 1M output tokens in us-east-1.
Which AWS regions support Llama 3.1 70B Instruct?
3 of 16 regions: us-east-1, us-east-2, us-west-2. Regions marked profile-only need a cross-region inference profile ID such as us.meta.llama3-1-70b-instruct-v1:0 rather than the bare model ID.
What is the context window of Llama 3.1 70B Instruct?
128,000 tokens (128K).
What is the model ID for Llama 3.1 70B Instruct on Bedrock?
meta.llama3-1-70b-instruct-v1:0. In regions where it is only reachable through a cross-region inference profile, use a prefixed ID instead: us.meta.llama3-1-70b-instruct-v1:0.

Other Meta models