Open-sourcing lithos-metal: Generating Megakernels for Apple Silicon
· 11 min read
We are excited to release lithos-metal, a fully open-source inference engine that generates Metal megakernels for Apple silicon.
lithos-metal brings ultra-fast inference from the cloud to the edge, serving Qwen3.8-27B at more than 200 tokens per second per user on M5 Max chips. For interactive applications such as local coding assistants, faster generation shortens response times and the feedback loop between the model and the user.
The design combines two complementary techniques. Megakernels coordinate computation across operator boundaries within a model pass. DSpark speculative decoding lets the target model verify several proposed tokens together, reusing loaded weights across the block. Efficient execution and efficient use of each target pass are both necessary for low-latency local inference.
What a megakernel changes
Conventional GPU execution assigns individual operators or small fused groups to separate kernel dispatches. Intermediate tensors connect the kernels, and synchronization enforces their dependencies. These boundaries constrain how worker-local results can be reused and how tasks from different operators can be interleaved.
The task-based execution model studied in Mirage Persistent Kernel (MPK) moves this scheduling into the GPU program. The compiler decomposes operators into tasks, such as matrix-output tiles, attention partitions, and recurrent-state slices, and records their dependencies. Persistent workers execute these tasks within one dispatch. In lithos-metal, a worker is a Metal threadgroup that processes a sequence of tasks until the fused region completes.
A task graph exposes independent work that can be scheduled together and makes the lifetime of intermediate results explicit. Values used within one worker may remain in registers or threadgroup memory. Results shared between workers still require device memory and synchronization before dependent tasks can consume them.
lithos-metal generates specialized Metal code for the tasks and selects a schedule for the target GPU. The compiled program combines mixer megakernels with conventional kernels and is replayed through a pre-encoded Metal indirect command buffer. GPU-resident state tracks token positions, active rows, and accepted tokens, with minimal CPU involvement.
Why megakernels fit Apple silicon
Data reuse. Apple’s measurements of LLM decoding on M5 show that memory bandwidth can dominate generation time. This makes reducing data movement an important complement to faster matrix arithmetic. A megakernel can retain intermediate values within the worker that consumes them, avoiding selected writes, rereads, and layout conversions. Normalization statistics, attention accumulators, and recurrent-state slices are useful examples: their reuse spans several computation steps, giving the compiler an opportunity to reduce traffic across the whole fused region.
Matrix acceleration. On M5, GPU neural accelerators provide matrix operations through Metal Performance Primitives. Metal supports inline tensor operations inside a shader, so the same workers can execute matrix tiles and surrounding normalization, gating, and reduction work. Cooperative tensors distribute matrix values across participating threads, allowing subsequent operations to consume them within the shader. This gives a megakernel a direct way to combine accelerated matrix arithmetic with the custom operations that make up a mixer.
Dynamic memory. Apple GPUs also provide dynamic on-chip memory allocation, introduced with the M3 generation. Register storage adjusts to the needs of the executing code, and on-chip capacity is shared across register, threadgroup, and buffer data. This is a useful match for a program whose projection, reduction, and elementwise stages need different amounts of temporary storage at different times. To use that flexibility effectively, lithos-metal tunes worker count, tile geometry, and intermediate lifetimes together, balancing local reuse against the storage needed to keep multiple workers active.
Layer-wise fusion. Instead of a whole-model megakernel, lithos-metal uses a layer-wise approach: the mixer in a GDN layer or a full-attention layer is converted into a single megakernel, while the layer’s MLP remains separate. This boundary captures reuse among tightly coupled operations and keeps dispatches bounded. Our M3 Pro and M5 Pro probes found that in-kernel global barriers were no cheaper than dispatch boundaries, and long dispatches could delay other GPU clients. Pre-encoded command replay connects the layer-wise megakernels and remaining kernels with little CPU work.
Mixer megakernel design
We use Qwen3.8-27B as a case study to demonstrate lithos-metal’s megakernel design. The model uses two forms of token mixing: Gated DeltaNet and full attention. A mixer combines information across tokens through projections, a recurrent or attention core, and an output projection. The examples below fuse this region while retaining separate kernels for the layer’s feed-forward network (MLP).
Gated DeltaNet (GDN) summarizes token history in a recurrent state whose size is independent of context length. Its main scheduling challenge is to preserve sequential updates across tokens while exposing parallel work within each update. The mixer projects queries, keys, and values (Q/K/V), recurrence controls A/B, and an output gate Z, then applies a causal convolution, normalization, the recurrence, gated normalization, and an output projection.

For one head, a simplified recurrence is:
S̄ₜ = αₜ Sₜ₋₁
δₜ = βₜ (vₜ − S̄ₜᵀ kₜ)
Sₜ = S̄ₜ + kₜ δₜᵀ
oₜ = Sₜᵀ qₜSₜ is the FP32 key-by-value state matrix; αₜ controls decay and βₜ controls update magnitude. The query and key vectors include normalization, with query scaling absorbed into qₜ. Each token depends on the previous state, but different value columns can be updated independently. We therefore assign each task a slice of state columns and retain that slice in registers across a block of input tokens. This avoids reading and writing the full recurrent state at every token step.
The task graph captures the remaining dependencies. QKV and A/B projections supply the recurrence, while the Z projection can be computed independently. Gated normalization waits for both the recurrent output and Z, then combines the output slices for each head. Convolution history and recurrent state are explicit outputs, preserving the state associated with the committed token prefix.
The selected M5 Max configuration uses 80 logical workers, eight SIMD groups per worker, and eight-column recurrence slices. Worker count is a scheduling parameter, independent of the physical core count. Device fences and bounded barriers make intermediate results visible before dependent tasks execute.
The compiler also accounts for work already performed by adjacent kernels. When a preceding residual kernel produces the normalized projection layout, the mixer reuses it. This removes redundant transformation work at the fusion boundary.
Full attention reads a key/value (K/V) cache that grows with context length. Longer contexts expose more parallel work, but also increase memory traffic and the number of partial results to combine. The mixer includes QKV and gate projections, Q/K normalization and rotary position embeddings (RoPE), causal attention, output gating, and an output projection with residual addition.

We partition attention over query tiles and key ranges. Small partitions provide more independent tasks, but each produces a partial result that must be stored and merged. To balance these costs, a task processes several tiles of 16 or 32 keys and accumulates their contributions locally before writing one partial result.
An online softmax keeps this accumulation numerically stable without storing all attention scores. For one query row, a new score tile s and corresponding value matrix V update the running statistics as follows:
m′ = max(m, max(s))
a = exp(m − m′)
p = exp(s − m′)
l′ = a·l + sum(p)
o′ = a·o + pVHere m is the running maximum, l is the softmax denominator, and o is the unnormalized weighted-value sum. When the maximum changes, a rescales the previous accumulators to the new reference. These FP32 statistics are sufficient to merge partitions: the final reduction rescales them to a common maximum, sums their contributions, and divides the weighted-value sum by the denominator.
The query tile is reused throughout the local loop. Staging storage depends on the physical key tile size, so a task can process a larger partition without a proportionally larger temporary buffer. In the illustrated 8K configuration, each local partition contains up to 1,024 keys.
Workers receive initial assignments and then claim additional tasks from an atomic queue for each dependency phase. Attention partitions share a phase with the independent gate projection. A synchronization boundary ensures that both complete before reduction and output projection. This keeps task assignment flexible while preserving the model’s data dependencies.
Megakernelizing the DSpark head
DSpark combines a small, target-conditioned transformer that drafts a block of tokens with a sequential Markov head that models dependencies between proposals. The target verifies the block and commits the accepted prefix. When several proposals are accepted, the weight reads for one target pass are amortized over multiple output tokens. Our configuration uses a five-layer drafter and seven proposals, verified together with an anchor token: the most recently committed token.
The benefit depends on keeping drafting cheaper than the target work it saves. Each round adds projections, context attention, and sequential token corrections, with their own computation, memory traffic, and dispatch costs. The mixer is a useful fusion region because its projections, attention, and reductions share intermediate results. Megakernel generation lets us coordinate those stages within one dispatch.

The same task compiler fuses block QKV projection, matrix attention, partial-result reduction, output projection, residual addition, and normalization for the next stage. Draft attention reads cached target context and newly injected target features alongside the proposal block, whose positions attend to one another bidirectionally. Task queues distribute projection and attention tiles across workers; local softmax accumulation reduces intermediate writes. The resulting schedule applies the same task and dependency model used by the target mixers.
Results
The target-compute projection below shows that lithos-metal achieves approximately 1.8–2× the throughput of MLX and Ollama at short contexts, and more than 2× that of vLLM-Metal. Throughput remains high as context length increases, reaching approximately 127 projected tokens per second at 32K context, with a widening advantage over all three backends.

Join the lithos-metal community
lithos-metal is fully open source under the Apache 2.0 license. We invite researchers, developers, and Apple silicon users to help make fast local inference available across more models and devices. Explore the code on GitHub, try it on your Mac, and share what you learn from your workloads.
We welcome contributions to kernel optimization, model support, benchmarks across Apple GPUs, correctness tests, and documentation. Testing on different Macs helps us extend the engine beyond the configurations studied here. Start with our contribution guide, or open an issue to report a bug, share reproducible measurements, or propose an improvement.
To try lithos-metal, follow the README’s quick start and begin with a short prompt from a coding or writing task you know well. Compare the response time and output with your current setup, then repeat with a longer conversation. When sharing results, include your Mac model, context length, and sampling settings so others can reproduce them.