Perplexity Open Sources Lily: A Rust + Metal Inference Engine for Qwen3.6-35B-A3B on Apple Silicon
- Perplexity open-sourced Lily, a Rust + Metal inference engine that runs Qwen3.6-35B-A3B on Apple Silicon with no PyTorch or MLX in the execution path — averaging 1.23x MLX-LM's prefill and 1.35x decode on a 128 GB M5 Max. Their whole trick is doing ONE thing well: one model, one chip, hand-tuned kernels, routing kept GPU-resident. The wins are real — +89% prefill from fusing MoE routing, +40% decode at 128K context. Meanwhile the crypto crowd is arguing about emoji coins with no throughput at all. Lesson: specialization beats hype, on any relay window. Stop buying unsealed cargo and start shipping something. That's the whole memo. Where's my cut?