The fastest agentic inference.

Three tiers. Every tier serves the same weights at full model quality.

Parameters
2.8T
Context
1M
Released
July 2026

Kimi K3 pricing per 1M tokens

Base
The floor. It's high.
Early access discount
$15.00
$12.00/ 1M output
$2.40/ 1M input
$0.24/ 1M cached
Latency
80+tok/s/user
Fast
You'll stop reading along.
Early access discount
$22.50
$20.00/ 1M output
$4.00/ 1M input
$0.40/ 1M cached
Latency
180+tok/s/user
Ultra
We stopped negotiating with transistors.
$X.XX/ 1M output
Priced per workload
Latency
250–1000tok/s/user

In every tier

Base, Fast, and Ultra serve the same open weights.

Full model quality

Full precision open weights, not quantized, with the full context window on every tier.

Drop-in endpoints

OpenAI and Anthropic compatible, streaming and tool calling included. Switch tier or model per request.

Predictable spend

Live usage in the dashboard, hard caps when you need them, and alerts before anything surprises you.

Questions