cuDNN Frontend is NVIDIA's modern, open-source entry point to the cuDNN library and a growing collection of high-performance open-source kernels.
-
Updated
Aug 15, 2026 - Python
cuDNN Frontend is NVIDIA's modern, open-source entry point to the cuDNN library and a growing collection of high-performance open-source kernels.
High-performance grouped matrix multiplication for fine-grained and ultra-fine-grained MoE workloads.
Workload-aware persistent expert-tile scheduling for irregular MoE inference kernels on NVIDIA Blackwell GPUs
Hand-built MoE expert parallelism: routing-skew capture, fused grouped-GEMM Triton kernels, 2-all2all dispatch/combine, overlap ablations — proving by counter-example why DeepEP exists.
Add a description, image, and links to the grouped-gemm topic page so that developers can more easily learn about it.
To associate your repository with the grouped-gemm topic, visit your repo's landing page and select "manage topics."