Blazing fast inference on standard GPUs
Optimize for the speed each agent experiences.
Peak tokens per second per user measured on Kimi K3, a 2.8-trillion-parameter model, on a single 8×B300 node, August 2026. The without-Lithos range is the fastest major providers serving the same model.
Tokens per second, per user
Per user speed determines how quickly agents complete real work. Aggregate throughput hides what each agent experiences.
Low latency across the loop
Plan, call tools, observe, retry. Lithos keeps the entire agent loop moving across thousands of steps.
Native precision, full quality
The speed comes from the serving system, not from degrading the model. You get the same weights and quality, only faster.
Get Lithos
Lithos API
A hosted OpenAI and Anthropic compatible endpoint. Point your agents at a new base URL and they just run faster.
- Any agent harness, just one line changed: the API URL
- Serves Kimi K3, GLM 5.3, Qwen 3.8, DeepSeek V4, GPT OSS, and many more
- Engine dynamically tuned to your agents
- No infrastructure to provision or operate
- Start in minutes with an API key
On-Prem
The Lithos Engine, deployed on your own standard GPUs, inside your network.
- Runs on the GPU hardware you already have
- Serves Kimi K3, GLM 5.3, Qwen 3.8, DeepSeek V4, GPT OSS, and many more
- Engine dynamically tuned to your agents
- Your prompts and data never leave your VPC
- Same speed and full model quality as the hosted API