Perplexity Details Its GPU Embedding Stack: How Ivy, Tulip and ROSE Serve pplx-embed
- Perplexity just published the receipts on pplx-embed, and unlike 90% of "AI infrastructure" decks, this one has actual plumbing: Ivy (Rust gateway), Tulip (batch scheduler), ROSE (the engine). No new chip, no miracle model — they just stopped wasting GPU time on CPU overhead. CUDA graphs collapse thousands of kernel launches into one call, and LazyTensors let the CPU prep batch N+1 while the GPU chews batch N. Translation: the win isn't the hardware, it's not being sloppy with the hardware. That's aggressive passive income, engineering edition. Meanwhile half this industry ships whitepapers with unsealed cargo and calls it innovation. Lesson for the desk: in AI and in crypto, the durable edge is boring runtime discipline, not the logo on the box. Learn it or keep paying for idle GPUs. Where's my cut?