At Lithos we hunt milliseconds through the whole stack, from GPU kernels to schedulers to serving. If you think our current speed is not fast enough, let's talk.
We are hiring across machine learning, inference optimization, and core systems, along with non-engineering roles across go-to-market. If you do not see the right role but think you belong here, reach out at jobs@lithosai.com
Benefits: Equity, generous health, dental, and vision benefits, plus 401(k) company match.
Visa sponsorship: We sponsor visas on a case-by-case basis.
Inference Optimization Engineer
San Francisco, CA or Pittsburgh, PA
Make frontier model inference faster end-to-end. You will work across GPU kernels, compilers, runtimes, batching, scheduling, and serving to turn systems research into production performance.
BS, MS, or PhD in Computer Science or related field.
Apply →Machine Learning Engineer
San Francisco, CA or Pittsburgh, PA
Make the models themselves faster to serve. You will work on speculative decoding, attention and KV cache optimizations, and model-level techniques that raise tokens per second while preserving full model quality.
BS, MS, or PhD in Computer Science or related field.
Apply →Generalist Systems Engineer
San Francisco, CA or Pittsburgh, PA
Build the distributed systems behind fast, reliable model serving. You will work across core infrastructure, resource management, fault tolerance, observability, and APIs as a generalist on the systems that run Lithos at scale.
BS, MS, or PhD in Computer Science or related field.
Apply →Developer Relations Engineer
Remote or San Francisco, CA or Pittsburgh, PA
Help developers build faster agents with Lithos. You will create technical content and examples, work directly with users, represent their needs to the engineering team, and grow a community around high performance model inference.
BS, MS, or PhD in Computer Science or related field.
Apply →