Optimizing 405B Parameter LLM Inference on Liquid-Cooled B200 Clusters
How we redesigned tensor parallelism communication over NVLink 5.0 and implemented custom FP8 GEMM kernels with low-rank factor decomposition, achieving 68 tokens/sec per stream on 405B parameter models.
Read Article →