The fastest agentic inference on the planet.

Full model quality. Standard GPUs.

Blazing fast inference on standard GPUs

Optimize for the speed each agent experiences.

HARNESSClaude Code•Codex•Your harness
MODELKimi K3•DeepSeek V4.1 Flash•GLM-5.3•Qwen 3.8•Gemma 4•GPT OSS•Nemotron•Your model

Without Lithos

50–150 tokens/second
pool.rs
Major Providers (Fastest Mode)
Speed
– tok/s
Generated
– tokens
Time
– s

With Lithos

800 tokens/second up to 16× faster
pool.rs
Lithos Engine
Speed
– tok/s
Generated
– tokens
Time
– s

Peak tokens per second per user measured on Kimi K3, a 2.8-trillion-parameter model, on a single 8×B300 node, August 2026. The without-Lithos range is the fastest major providers serving the same model.

  • Tokens per second, per user

    Per user speed determines how quickly agents complete real work. Aggregate throughput hides what each agent experiences.

  • Low latency across the loop

    Plan, call tools, observe, retry. Lithos keeps the entire agent loop moving across thousands of steps.

  • Native precision, full quality

    The speed comes from the serving system, not from degrading the model. You get the same weights and quality, only faster.

Get Lithos

Lithos API

A hosted OpenAI compatible endpoint. Point your agents at a new base URL and they just run faster.

  • Any agent harness, just one line changed: the API URL
  • Serves Kimi K3, DeepSeek V4.1 Flash, GLM-5.3, Qwen 3.8, GPT OSS, and many more
  • Engine dynamically tuned to your agents
  • No infrastructure to provision or operate
  • Start in minutes with an API key

On-Prem

The Lithos Engine, deployed on your own standard GPUs, inside your network.

  • Runs on the GPU hardware you already have
  • Serves Kimi K3, DeepSeek V4.1 Flash, GLM-5.3, Qwen 3.8, GPT OSS, and many more
  • Engine dynamically tuned to your agents
  • Your prompts and data never leave your VPC
  • Same speed and full model quality as the hosted API

See what your agents can do when the loop moves faster