Inside NVIDIA’s cuDNN Graph API: Fusion, Autotuning, and Plan Reuse with cuDNN Frontend
- NVIDIA's cuDNN Frontend Graph API tutorial dropped, and unlike your average moonboy whitepaper, this thing actually has substance under the hull. You describe computation as a graph of ops, let cuDNN pick the engine, then — here's the fun part — take the wheel yourself. Fused conv-bias-ReLU, autotuning across engine configs, FP8-style epilogues, scaled dot-product attention, dynamic shapes, CUDA graph captures. All validated against PyTorch on a single Colab GPU, because in this business you check the manifest before you sign the PoD seal. Nobody's promising 400% APY here; they're promising correctness and measured latency, which is somehow the more radical pitch in 2025. Framework abstractions are comfy until your stack-eye catches them eating your FLOPs. Learn to fuse kernels below the framework and you stop renting efficiency from middlemen. Aggressive passive income, GPU edition. Where's my cut? Probably buried in the autotuner's engine configs — go dig it out yourself.