A faster slice of Pi.
Qwen3.8 in Pi, 3–4× faster than other libraries.* Pi stays Pi. You just wait a lot less.
Windows + NVIDIA · Apple Silicon Mac (* vs llama.cpp)
THE WHOLE PI
The whole damn pie.
Same weights. More slices.
Measured inside Pi on a real coding task, not a synthetic benchmark.
193tok/s
Lemon Pi
50tok/s
llama.cpp
3–4×
Faster · same model, same GPU
* NVIDIA RTX 3090 (24 GB) · Qwen3.8-27B · one Pi coding task · vs llama.cpp CUDA on the same model and GPU.
Baked for your GPU.
Speed comes from how the weights, kernels and execution graph fit together, so we bake them together.
- 1Mix
Three ingredients.
Qwen3.8's weights, the kernels that read them, and the execution graph that schedules them.
- 2Bake
Into the oven.
Kernels are compiled for your GPU, and weights are packed the way the kernels read them.
- 3Serve
Out comes a .uma.
One file with everything baked in. Warp opens it and picks the fastest path for your GPU.
No half-baked answers.
Is it open source?
The Lemon Pi plugin is Apache-2.0. The Warp runtime it ships with is proprietary.
What hardware do I need?
An NVIDIA GPU on Windows, or an Apple Silicon Mac. Qwen3.8-27B fits a 24 GB card like the RTX 3090. Linux is in the oven.
What about llama.cpp with speculative decoding?
Still about 2.5× faster. With its own speculative decoding, llama.cpp does 76 tok/s on the same model and GPU; Lemon Pi does 193.
Does it change the model's output?
No. Every speculative token is verified before it's kept, and every build's output is checked by an oracle.
Why is the first load slow?
A 27B model takes a couple of minutes to load. Pi preheats it in the background as soon as it opens, then keeps it warm between sessions.
What about other models?
Qwen3.8 first. More models will show up on the menu as they're baked.
Get a slice.
1.Get PiSkip if you have it.
2.Add Lemon PiAdds Warp to Pi.
3.Grab Qwen3.8 and goInside Pi. Checks your GPU before it downloads.
Windows + NVIDIA or an Apple Silicon Mac