Extends expert-parallelism communication with a handful of novel optimizations — adaptive compute-resource allocation, topology-aware routing, and a mixed-precision transfer pipeline — validated on multi-GPU hardware.
Back to projects
Jun 14, 2026
1 min read
DeepEP-X
A CUDA engine extending expert-parallel GPU communication with adaptive resource allocation and smarter routing.