The fastest agentic inference.
Every tier serves the same weights at full model quality.
- Parameters
- 2.8T
- Context
- 1M
- Released
- July 2026
- Parameters
- 552B
- Context
- 1M
- Released
- September 2026
- Parameters
- 753B
- Context
- 1M
- Released
- August 2026
- Parameters
- 321B
- Context
- 1M
- Released
- August 2026
Kimi K3 pricing per 1M tokens
Base
The floor. It's high.
Early access discount
$12.00 / 1M output
$2.40 / 1M input
$0.24 / 1M cached
Latency
80+ tok/s/user
Fast
You'll stop reading along.
Early access discount
$20.00 / 1M output
$4.00 / 1M input
$0.40 / 1M cached
Latency
180+ tok/s/user
Ultra
We stopped negotiating with transistors.
Early access discount
$28.00 / 1M output
$5.60 / 1M input
$0.56 / 1M cached
Latency
250–1000 tok/s/user
DeepSeek V4.1 Flash pricing per 1M tokens
Base
The floor. It's high.
Early access discount
$0.60 / 1M output
$0.15 / 1M input
$0.003 / 1M cached
Latency
250+ tok/s/user
Fast
You'll stop reading along.
Early access discount
$1.00 / 1M output
$0.25 / 1M input
$0.005 / 1M cached
Latency
350+ tok/s/user
Ultra
We stopped negotiating with transistors.
Early access discount
$1.40 / 1M output
$0.35 / 1M input
$0.007 / 1M cached
Latency
450+ tok/s/user
GLM-5.3 pricing per 1M tokens
Base
The floor. It's high.
Early access discount
$3.30 / 1M output
$1.05 / 1M input
$0.195 / 1M cached
Latency
200+ tok/s/user
GLM-5.3-Flash pricing per 1M tokens
Base
The floor. It's high.
$0.50 / 1M output
$0.15 / 1M input
$0.03 / 1M cached
Latency
250+ tok/s/user
Qwen 3.8 availability
Qwen 3.8
Coming soon to the Lithos engine.
Tell us what you are building on Qwen 3.8 and we will let you know the moment it is serving traffic.
Gemma 4 availability
Gemma 4
Coming soon to the Lithos engine.
Tell us what you are building on Gemma 4 and we will let you know the moment it is serving traffic.
GPT OSS availability
GPT OSS
Coming soon to the Lithos engine.
Tell us what you are building on GPT OSS and we will let you know the moment it is serving traffic.
Bring your own model
Bring Your Own Model (BYOM)
Any open model, or bring your own.
$X.XX / 1M output
Priced per workload
Latency
Sized by your model
In every tier
Base, Fast, and Ultra serve the same open weights.
Full model quality
Full precision open weights, not quantized, with the full context window on every tier.
Drop-in endpoints
OpenAI compatible, streaming and tool calling included. Switch tier or model per request.
Predictable spend
Live usage in the dashboard, hard caps when you need them, and alerts before anything surprises you.