← All posts

Open-sourcing lithos-metal: Generating Megakernels for Apple Silicon

· 11 min read

We are excited to release lithos-metal, a fully open-source inference engine that generates Metal megakernels for Apple silicon.

lithos-metal brings ultra-fast inference from the cloud to the edge, serving Qwen3.8-27B at more than 200 tokens per second per user on M5 Max chips. For interactive applications such as local coding assistants, faster generation shortens response times and the feedback loop between the model and the user.

Watch lithos-metal run Qwen3.8-27B with DSpark locally on Apple silicon.

The design combines two complementary techniques. Megakernels coordinate computation across operator boundaries within a model pass. DSpark speculative decoding lets the target model verify several proposed tokens together, reusing loaded weights across the block. Efficient execution and efficient use of each target pass are both necessary for low-latency local inference.

What a megakernel changes

Conventional GPU execution assigns individual operators or small fused groups to separate kernel dispatches. Intermediate tensors connect the kernels, and synchronization enforces their dependencies. These boundaries constrain how worker-local results can be reused and how tasks from different operators can be interleaved.

The task-based execution model studied in Mirage Persistent Kernel (MPK) moves this scheduling into the GPU program. The compiler decomposes operators into tasks, such as matrix-output tiles, attention partitions, and recurrent-state slices, and records their dependencies. Persistent workers execute these tasks within one dispatch. In lithos-metal, a worker is a Metal threadgroup that processes a sequence of tasks until the fused region completes.

Separate kernels3 dispatches
Separate operator dispatchesThree numbered launch arrows enter separate kernels for input projection, attention, and output projection. Each kernel contains parallel tasks. Vertical arrows show dependencies, not elapsed time.1Input projectionP0P1P22AttentionA0A1A23Output projectionO0O1O2
Megakernel1 dispatch
GPU task dispatch to persistent workersOne launch enters a megakernel. Persistent workers claim ready tasks from a shared queue and execute projection, attention, and output tasks as dependencies are satisfied. Colored task blocks correspond to the same operators on the left. Horizontal joins synchronize dependent phases. The lanes illustrate task assignment, not physical cores or a measured timeline.Ready tasksP0P1P2Worker 0P0A0O0Worker 1P1A1O1Worker 2P2A2O2
Persistent workers claim ready tasks across operator boundaries.

A task graph exposes independent work that can be scheduled together and makes the lifetime of intermediate results explicit. Values used within one worker may remain in registers or threadgroup memory. Results shared between workers still require device memory and synchronization before dependent tasks can consume them.

lithos-metal generates specialized Metal code for the tasks and selects a schedule for the target GPU. The compiled program combines mixer megakernels with conventional kernels and is replayed through a pre-encoded Metal indirect command buffer. GPU-resident state tracks token positions, active rows, and accepted tokens, with minimal CPU involvement.

Why megakernels fit Apple silicon

Data reuse. Apple’s measurements of LLM decoding on M5 show that memory bandwidth can dominate generation time. This makes reducing data movement an important complement to faster matrix arithmetic. A megakernel can retain intermediate values within the worker that consumes them, avoiding selected writes, rereads, and layout conversions. Normalization statistics, attention accumulators, and recurrent-state slices are useful examples: their reuse spans several computation steps, giving the compiler an opportunity to reduce traffic across the whole fused region.

Matrix acceleration. On M5, GPU neural accelerators provide matrix operations through Metal Performance Primitives. Metal supports inline tensor operations inside a shader, so the same workers can execute matrix tiles and surrounding normalization, gating, and reduction work. Cooperative tensors distribute matrix values across participating threads, allowing subsequent operations to consume them within the shader. This gives a megakernel a direct way to combine accelerated matrix arithmetic with the custom operations that make up a mixer.

Dynamic memory. Apple GPUs also provide dynamic on-chip memory allocation, introduced with the M3 generation. Register storage adjusts to the needs of the executing code, and on-chip capacity is shared across register, threadgroup, and buffer data. This is a useful match for a program whose projection, reduction, and elementwise stages need different amounts of temporary storage at different times. To use that flexibility effectively, lithos-metal tunes worker count, tile geometry, and intermediate lifetimes together, balancing local reuse against the storage needed to keep multiple workers active.

Layer-wise fusion. Instead of a whole-model megakernel, lithos-metal uses a layer-wise approach: the mixer in a GDN layer or a full-attention layer is converted into a single megakernel, while the layer’s MLP remains separate. This boundary captures reuse among tightly coupled operations and keeps dispatches bounded. Our M3 Pro and M5 Pro probes found that in-kernel global barriers were no cheaper than dispatch boundaries, and long dispatches could delay other GPU clients. Pre-encoded command replay connects the layer-wise megakernels and remaining kernels with little CPU work.

Mixer megakernel design

We use Qwen3.8-27B as a case study to demonstrate lithos-metal’s megakernel design. The model uses two forms of token mixing: Gated DeltaNet and full attention. A mixer combines information across tokens through projections, a recurrent or attention core, and an output projection. The examples below fuse this region while retaining separate kernels for the layer’s feed-forward network (MLP).

Gated DeltaNet (GDN) summarizes token history in a recurrent state whose size is independent of context length. Its main scheduling challenge is to preserve sequential updates across tokens while exposing parallel work within each update. The mixer projects queries, keys, and values (Q/K/V), recurrence controls A/B, and an output gate Z, then applies a causal convolution, normalization, the recurrence, gated normalization, and an output projection.

GDN computation graph and proposed task schedule for 80 persistent worker groups
GDN computation and task dependencies. The right panel illustrates a proposed balanced task pool; the measured configuration uses staged scheduling. Open the figure to zoom.

For one head, a simplified recurrence is:

S̄ₜ = αₜ Sₜ₋₁
δₜ = βₜ (vₜ − S̄ₜᵀ kₜ)
Sₜ = S̄ₜ + kₜ δₜᵀ
oₜ = Sₜᵀ qₜ

Sₜ is the FP32 key-by-value state matrix; αₜ controls decay and βₜ controls update magnitude. The query and key vectors include normalization, with query scaling absorbed into qₜ. Each token depends on the previous state, but different value columns can be updated independently. We therefore assign each task a slice of state columns and retain that slice in registers across a block of input tokens. This avoids reading and writing the full recurrent state at every token step.

The task graph captures the remaining dependencies. QKV and A/B projections supply the recurrence, while the Z projection can be computed independently. Gated normalization waits for both the recurrent output and Z, then combines the output slices for each head. Convolution history and recurrent state are explicit outputs, preserving the state associated with the committed token prefix.

The selected M5 Max configuration uses 80 logical workers, eight SIMD groups per worker, and eight-column recurrence slices. Worker count is a scheduling parameter, independent of the physical core count. Device fences and bounded barriers make intermediate results visible before dependent tasks execute.

The compiler also accounts for work already performed by adjacent kernels. When a preceding residual kernel produces the normalized projection layout, the mixer reuses it. This removes redundant transformation work at the fusion boundary.

Megakernelizing the DSpark head

DSpark combines a small, target-conditioned transformer that drafts a block of tokens with a sequential Markov head that models dependencies between proposals. The target verifies the block and commits the accepted prefix. When several proposals are accepted, the weight reads for one target pass are amortized over multiple output tokens. Our configuration uses a five-layer drafter and seven proposals, verified together with an anchor token: the most recently committed token.

The benefit depends on keeping drafting cheaper than the target work it saves. Each round adds projections, context attention, and sequential token corrections, with their own computation, memory traffic, and dispatch costs. The mixer is a useful fusion region because its projections, attention, and reductions share intermediate results. Megakernel generation lets us coordinate those stages within one dispatch.

DSpark block pipeline, per-layer mixer fusion boundary, and seven-position sequential Markov chain
The orange region denotes one mixer megakernel per draft layer in the BF16 configuration selected at 128 and 32K context. Context-K/V injection and the two MLP kernels remain separate. Open the figure to zoom.

The same task compiler fuses block QKV projection, matrix attention, partial-result reduction, output projection, residual addition, and normalization for the next stage. Draft attention reads cached target context and newly injected target features alongside the proposal block, whose positions attend to one another bidirectionally. Task queues distribute projection and attention tiles across workers; local softmax accumulation reduces intermediate writes. The resulting schedule applies the same task and dependency model used by the target mixers.

Results

The target-compute projection below shows that lithos-metal achieves approximately 1.8–2× the throughput of MLX and Ollama at short contexts, and more than 2× that of vLLM-Metal. Throughput remains high as context length increases, reaching approximately 127 projected tokens per second at 32K context, with a widening advantage over all three backends.

Conditional throughput projection for Lithos Metal, MLX, Ollama, and vLLM-Metal from 128 to 32K context
Qwen3.8-27B-NVFP4 on a 40-core M5 Max with 48 GB.

Join the lithos-metal community

lithos-metal is fully open source under the Apache 2.0 license. We invite researchers, developers, and Apple silicon users to help make fast local inference available across more models and devices. Explore the code on GitHub, try it on your Mac, and share what you learn from your workloads.

We welcome contributions to kernel optimization, model support, benchmarks across Apple GPUs, correctness tests, and documentation. Testing on different Macs helps us extend the engine beyond the configurations studied here. Start with our contribution guide, or open an issue to report a bug, share reproducible measurements, or propose an improvement.

To try lithos-metal, follow the README’s quick start and begin with a short prompt from a coding or writing task you know well. Compare the response time and output with your current setup, then repeat with a longer conversation. When sharing results, include your Mac model, context length, and sampling settings so others can reproduce them.