Embed v4
Cohere · Active · available in 16 of 16 regions
- Context window
- 128K
- 128,000 tokens
- Max output
- —
- Input / 1M
- —
- Output / 1M
- —
Model & inference profile IDs
Base model ID — direct on-demand invoke
Cross-region inference profiles (3)
Regions marked Inference profile only require a profile ID rather than the base model ID — invoking the base ID there returns a validation error. Regional (us., eu.) profiles carry a 10% premium over global.
Pricing by region
USD per 1M tokensNo published token pricing
The AWS Price List API publishes no token SKU for this model, and it is not in our curated overlay. This is common for embedding, image and video models, which bill per request or per second rather than per token.
Check the AWS pricing page →Region availability
TPM / RPM are default account quotas from AWS Service Quotas where a model-specific limit is published. They are per-account defaults and adjustable on request.
Common questions
- Which AWS regions support Embed v4?
- 16 of 16 regions: us-east-1, us-east-2, us-west-2, ca-central-1, eu-west-1, eu-west-2, eu-west-3, eu-central-1, eu-north-1, ap-southeast-1, ap-southeast-2, ap-northeast-1, ap-northeast-2, ap-northeast-3, ap-south-1, sa-east-1. Regions marked profile-only need a cross-region inference profile ID such as eu.cohere.embed-v4:0 rather than the bare model ID.
- What is the context window of Embed v4?
- 128,000 tokens (128K).
- What is the model ID for Embed v4 on Bedrock?
- cohere.embed-v4:0. In regions where it is only reachable through a cross-region inference profile, use a prefixed ID instead: eu.cohere.embed-v4:0, global.cohere.embed-v4:0, us.cohere.embed-v4:0.