Bedrock models with batch inference
These Amazon Bedrock models support batch inference, which processes a submitted job asynchronously at roughly half the on-demand token rate. If the work can wait hours rather than seconds, this is the simplest halving of a token bill available on the platform.
56 of the models on Bedrock qualify today, from $0.035 per 1M input tokens. Rebuilt daily from the AWS catalogue, so this list does not go stale.
What teams get wrong about this
Batch is not a latency setting you can flip on an existing endpoint — it is a different API with a different shape. You submit a file of records to S3 and collect results later, so it fits backfills, nightly enrichment and evaluation runs, and does not fit anything a user is waiting on. It also does not combine with prompt caching.
All 56 models, cheapest first
| Model | Input /1M | Batch input |
|---|---|---|
| Nova Micro amazon.nova-micro-v1:0 | $0.035 | $0.02 |
| Gemma 3 4B IT google.gemma-3-4b-it | $0.040 | $0.02 |
| Voxtral Mini 3B 2507 mistral.voxtral-mini-3b-2507 | $0.040 | $0.02 |
| Nova Lite amazon.nova-lite-v1:0 | $0.060 | $0.03 |
| Nemotron Nano 3 30B nvidia.nemotron-nano-3-30b | $0.060 | $0.03 |
| NVIDIA Nemotron Nano 9B v2 nvidia.nemotron-nano-9b-v2 | $0.060 | $0.03 |
| GPT OSS Safeguard 20B openai.gpt-oss-safeguard-20b | $0.070 | $0.03 |
| gpt-oss-20b openai.gpt-oss-20b-1:0 | $0.070 | $0.03 |
| GLM 4.7 Flash zai.glm-4.7-flash | $0.070 | $0.03 |
| Gemma 3 12B IT google.gemma-3-12b-it | $0.090 | $0.05 |
| Ministral 3B mistral.ministral-3-3b-instruct | $0.100 | $0.05 |
| Qwen3 Next 80B A3B qwen.qwen3-next-80b-a3b | $0.140 | $0.07 |
| Ministral 3 8B mistral.ministral-3-8b-instruct | $0.150 | $0.07 |
| NVIDIA Nemotron 3 Super 120B A12B nvidia.nemotron-super-3-120b | $0.150 | $0.07 |
| GPT OSS Safeguard 120B openai.gpt-oss-safeguard-120b | $0.150 | $0.07 |
| gpt-oss-120b openai.gpt-oss-120b-1:0 | $0.150 | $0.07 |
| Qwen3 32B (dense) qwen.qwen3-32b-v1:0 | $0.150 | $0.07 |
| Qwen3-Coder-30B-A3B-Instruct qwen.qwen3-coder-30b-a3b-v1:0 | $0.150 | $0.07 |
| Writer Palmyra Vision 7B writer.palmyra-vision-7b | $0.150 | $0.07 |
| Llama 4 Scout 17B Instruct meta.llama4-scout-17b-instruct-v1:0 | $0.170 | $0.09 |
| Ministral 14B 3.0 mistral.ministral-3-14b-instruct | $0.200 | $0.10 |
| Llama 3.1 8B Instruct meta.llama3-1-8b-instruct-v1:0 | $0.220 | $0.11 |
| Qwen3 235B A22B 2507 qwen.qwen3-235b-a22b-2507-v1:0 | $0.220 | $0.11 |
| Gemma 3 27B PT google.gemma-3-27b-it | $0.230 | $0.12 |
| Llama 4 Maverick 17B Instruct meta.llama4-maverick-17b-instruct-v1:0 | $0.240 | $0.12 |
| Nova 2 Lite amazon.nova-2-lite-v1:0 | $0.300 | $0.15 |
| MiniMax M2 minimax.minimax-m2 | $0.300 | $0.15 |
| MiniMax M2.1 minimax.minimax-m2.1 | $0.300 | $0.15 |
| MiniMax M2.5 minimax.minimax-m2.5 | $0.300 | $0.15 |
| Qwen3 Coder 480B A35B Instruct qwen.qwen3-coder-480b-a35b-v1:0 | $0.450 | $0.23 |
| Magistral Small 2509 mistral.magistral-small-2509 | $0.500 | $0.25 |
| Mistral Large 3 mistral.mistral-large-3-675b-instruct | $0.500 | $0.25 |
| Qwen3 Coder Next qwen.qwen3-coder-next | $0.500 | $0.25 |
| Qwen3 VL 235B A22B qwen.qwen3-vl-235b-a22b | $0.530 | $0.26 |
| DeepSeek-V3.1 deepseek.v3-v1:0 | $0.580 | $0.29 |
| Kimi K2 Thinking moonshot.kimi-k2-thinking | $0.600 | $0.30 |
| Kimi K2.5 moonshotai.kimi-k2.5 | $0.600 | $0.30 |
| GLM 4.7 zai.glm-4.7 | $0.600 | $0.30 |
| DeepSeek V3.2 deepseek.v3.2 | $0.620 | $0.31 |
| Llama 3.1 70B Instruct meta.llama3-1-70b-instruct-v1:0 | $0.720 | $0.36 |
| Llama 3.3 70B Instruct meta.llama3-3-70b-instruct-v1:0 | $0.720 | $0.36 |
| Nova Pro amazon.nova-pro-v1:0 | $0.800 | $0.40 |
| Claude Haiku 4.5 anthropic.claude-haiku-4-5-20251001-v1:0 | $1.00 | $0.50 |
| Mistral Small (24.02) mistral.mistral-small-2402-v1:0 | $1.00 | $0.50 |
| GLM 5 zai.glm-5 | $1.00 | $0.50 |
| Mistral Large (24.07) mistral.mistral-large-2407-v1:0 | $2.00 | $1.50 |
| Nova Premier amazon.nova-premier-v1:0 | $2.50 | $1.25 |
| Claude Sonnet 4.5 anthropic.claude-sonnet-4-5-20250929-v1:0 | $3.00 | $1.50 |
| Claude Sonnet 4.6 anthropic.claude-sonnet-4-6 | $3.00 | $1.50 |
| Claude Sonnet 5 anthropic.claude-sonnet-5 | $3.00 | $1.50 |
| Claude Opus 4.5 anthropic.claude-opus-4-5-20251101-v1:0 | $5.00 | $2.50 |
| Claude Opus 4.6 anthropic.claude-opus-4-6-v1 | $5.00 | $2.50 |
| Claude Opus 4.7 anthropic.claude-opus-4-7 | $5.00 | $2.50 |
| Claude Opus 4.8 anthropic.claude-opus-4-8 | $5.00 | $2.50 |
| Claude Opus 5 anthropic.claude-opus-5 | $5.00 | $2.50 |
| Claude Fable 5 anthropic.claude-fable-5 | $10.00 | $5.00 |
Capability flags come from the AWS Bedrock model catalogue; prices from the Price List API. Region counts include cross-region inference profiles. Methodology.
Questions
How much cheaper is Bedrock batch inference?
Around 50% of the on-demand rate for both input and output on the models that support it. The exact figure per model and region is in the table on each model page.
How long does a Bedrock batch job take?
It is asynchronous with no latency guarantee — plan for hours, not minutes. That is the trade you are making for the discount.