The fastest agentic inference.
Three tiers. Every tier serves the same weights at full model quality.
- Parameters
- 2.8T
- Context
- 1M
- Released
- July 2026
Kimi K3 pricing per 1M tokens
GLM 5.3 availability
Tell us what you are building on GLM 5.3 and we will let you know the moment it is serving traffic.
Qwen 3.8 availability
Tell us what you are building on Qwen 3.8 and we will let you know the moment it is serving traffic.
DeepSeek V4 availability
Tell us what you are building on DeepSeek V4 and we will let you know the moment it is serving traffic.
Gemma 4 availability
Tell us what you are building on Gemma 4 and we will let you know the moment it is serving traffic.
GPT OSS availability
Tell us what you are building on GPT OSS and we will let you know the moment it is serving traffic.
In every tier
Base, Fast, and Ultra serve the same open weights.
Full model quality
Full precision open weights, not quantized, with the full context window on every tier.
Drop-in endpoints
OpenAI and Anthropic compatible, streaming and tool calling included. Switch tier or model per request.
Predictable spend
Live usage in the dashboard, hard caps when you need them, and alerts before anything surprises you.