Kimi K2 Thinking

Moonshot AI · Active · available in 7 of 31 regions

Batch inferenceStreaming Tool useVisionPrompt cachingEmbeddingsFine-tuning
Context window
Max output
Input / 1M
$0.600
us-east-1
Output / 1M
$2.50

What Kimi K2 Thinking is good for

editorial

Kimi K2 Thinking, and is the cheapest option from this provider at $0.600 per 1M input tokens. Callable in 7 of 31 regions.

Suits

  • Bulk jobs. Batch inference is available, typically at half the on-demand rate, for work that can wait — backfills, nightly enrichment, evaluation runs.
  • Cost. The cheapest Moonshot AI model here at $0.600 per 1M input tokens — the default choice for high-volume, low-difficulty work.

Think twice if

  • Agentic loops. No prompt caching published, so every step re-pays full price for the prompt prefix. Expensive for agents that resend a long context.
  • Function calling. No tool use, so it cannot drive an agent loop or call your APIs — it answers from what is in the prompt.
  • Documents with layout. Text only. Scanned PDFs, screenshots and charts need a vision model or an OCR step first.

This section is judgement, not data from AWS. It is composed from the capabilities, context window, price position and region coverage shown elsewhere on this page — so it stays in step with the daily snapshot rather than going stale.

Model & inference profile IDs

Base model ID — direct on-demand invoke

Pricing by region

USD per 1M tokens
Region Input Output Cache read Cache write Batch in Batch out Source
us-east-1 US East (N. Virginia) $0.600 $2.50 $0.300 $1.25 API
us-east-2 US East (Ohio) $0.600 $2.50 $0.300 $1.25 API
us-west-2 US West (Oregon) $0.600 $2.50 $0.300 $1.25 API
ap-northeast-1 Asia Pacific (Tokyo) $0.730 $3.03 $0.360 $1.52 API
ap-south-1 Asia Pacific (Mumbai) $0.710 $2.94 $0.360 $1.47 API
ap-southeast-2 Asia Pacific (Sydney) $0.618 $2.58 $0.309 $1.29 API
sa-east-1 South America (Sao Paulo) $0.730 $3.03 $0.360 $1.52 API

Region availability

TPM / RPM are default account quotas from AWS Service Quotas where a model-specific limit is published. They are per-account defaults and adjustable on request.

Where this model actually runs

Under a strict residency constraint, Kimi K2 Thinking can be served without the request leaving 5 of the 25 jurisdictions with a Bedrock region.

Kimi K2 Thinking appears in

Common questions

How much does Kimi K2 Thinking cost on Amazon Bedrock?
$0.600 per 1M input tokens and $2.50 per 1M output tokens in us-east-1.
Which AWS regions support Kimi K2 Thinking?
7 of 31 regions: us-east-1, us-east-2, us-west-2, ap-northeast-1, ap-south-1, ap-southeast-2, sa-east-1.
What is the model ID for Kimi K2 Thinking on Bedrock?
moonshot.kimi-k2-thinking.

Other Moonshot AI models