Llama 3.1 8B Instruct
Meta · Active · available in 3 of 16 regions
Tool useBatch inferenceStreaming VisionPrompt cachingEmbeddingsFine-tuning
- Context window
- 128K
- 128,000 tokens
- Max output
- —
- Input / 1M
- $0.220
- us-east-1
- Output / 1M
- $0.220
Model & inference profile IDs
Base model ID — direct on-demand invoke
Cross-region inference profiles (1)
Regions marked Inference profile only require a profile ID rather than the base model ID — invoking the base ID there returns a validation error. Regional (us., eu.) profiles carry a 10% premium over global.
Pricing by region
USD per 1M tokens| Region | Input | Output | Cache read | Cache write | Batch in | Batch out | Source |
|---|---|---|---|---|---|---|---|
| us-east-1 US East (N. Virginia) | $0.220 | $0.220 | — | — | $0.110 | $0.110 | API |
| us-east-2 US East (Ohio) | $0.220 | $0.220 | — | — | $0.110 | $0.110 | API |
| us-west-2 US West (Oregon) | $0.220 | $0.220 | — | — | $0.110 | $0.110 | API |
Region availability
Inference profile only us-east-1 Profile only Inference profile only us-east-2 Profile only On-demand us-west-2 On-demand Not available ca-central-1 Not available Not available eu-west-1 Not available Not available eu-west-2 Not available Not available eu-west-3 Not available Not available eu-central-1 Not available Not available eu-north-1 Not available Not available ap-southeast-1 Not available Not available ap-southeast-2 Not available Not available ap-northeast-1 Not available Not available ap-northeast-2 Not available Not available ap-northeast-3 Not available Not available ap-south-1 Not available Not available sa-east-1 Not available
TPM / RPM are default account quotas from AWS Service Quotas where a model-specific limit is published. They are per-account defaults and adjustable on request.
Common questions
- How much does Llama 3.1 8B Instruct cost on Amazon Bedrock?
- $0.220 per 1M input tokens and $0.220 per 1M output tokens in us-east-1.
- Which AWS regions support Llama 3.1 8B Instruct?
- 3 of 16 regions: us-east-1, us-east-2, us-west-2. Regions marked profile-only need a cross-region inference profile ID such as us.meta.llama3-1-8b-instruct-v1:0 rather than the bare model ID.
- What is the context window of Llama 3.1 8B Instruct?
- 128,000 tokens (128K).
- What is the model ID for Llama 3.1 8B Instruct on Bedrock?
- meta.llama3-1-8b-instruct-v1:0. In regions where it is only reachable through a cross-region inference profile, use a prefixed ID instead: us.meta.llama3-1-8b-instruct-v1:0.