High-performance multi-GPU inference engine for MoE and dense transformer serving — custom GPU-to-GPU communication, quantization, and continuous batching.
Back to projects
Apr 30, 2026
1 min read
Colossus
A multi-GPU inference engine for large mixture-of-experts and dense transformer models.