Cursor Open-Sources Mixture-of-Kittens (MoK): A Deterministic MoE Training Megakernel for GB300 NVL72 Racks

Kwon Crash

Published Aug 4, 2026, 10:03 PM UTC

Source: AISource
- Cursor open-sourced Mixture-of-Kittens (MoK), a deterministic MoE megakernel that’s 2.37x faster on GB300 NVL72 racks. It fuses communication and computation, using pull-based dispatch to slash signaling latency from 103 µs to 18 µs. The catch? It demands Blackwell SM100/SM103 GPUs, CUDA 13.0+, and PyTorch 2.10+. This isn’t for your 8-GPU garage setup; it’s for frontier labs with NVL72 capacity. While you’re busy chasing the "Altcoin of the week up 2000%," these guys are optimizing token routing at scale. It’s Apache-2.0, but the hardware floor is higher than most people’s credit score. If you don’t own a rack, this is just expensive theory. Professional solidarity: don't touch the robot, especially if you can't afford the electricity bill.