NVIDIA Mellanox 920-9B210-00FN-0D0 in Practice: Building an NDR Low-Latency Fabric for Trillion-Parameter MoE Models
August 27, 2026
NVIDIA Mellanox 920-9B210-00FN-0D0 in Practice: Building an NDR Low-Latency Fabric for Trillion-Parameter MoE Models
Background & The Challenge: When All-to-All Becomes the Training Bottleneck
A leading AI research organization recently scaled its large language model training cluster from 1,024 to 4,096 NVIDIA H100 GPUs, with the ambitious goal of reducing training time for a 1.8-trillion-parameter sparse MoE (Mixture of Experts) model from months to under two weeks. However, during initial validation, the team discovered that their existing 200Gb/s HDR network was experiencing severe congestion during the all-to-all communication phase—a characteristic pattern of MoE architectures—causing GPU utilization to plummet from an expected 85% to just 52%, far below the threshold needed to meet their training deadline.
Network analysis revealed that all-to-all collective communication accounted for over 40% of the MoE model's communication footprint, and as the number of experts scaled, the traffic pattern became increasingly fragmented and bursty. The existing HDR network, after triggering congestion control mechanisms, experienced dramatic tail latency increases, leaving a significant portion of GPUs idle and waiting. The organization urgently needed a network fabric capable of delivering deterministic sub-microsecond latency, lossless transport, and massive non-blocking scalability—and the NVIDIA Mellanox 920-9B210-00FN-0D0 was purpose-built for precisely this challenge.
Solution & Deployment: A MoE-Optimized NDR Leaf-Spine Architecture
The organization deployed a two-tier leaf-spine fabric built around the 920-9B210-00FN-0D0 InfiniBand switch OPN. In each compute rack, two 920-9B210-00FN-0D0 MQM9790-NS2F 400Gb/s NDR switches were deployed at the leaf layer, providing 400Gb/s connectivity to 64 dual-port H100 servers via OSFP-to-OSFP active optical cables. At the spine layer, 16 additional 920-9B210-00FN-0D0 switches interconnected all leaf switches, creating a fully non-blocking fat-tree fabric with an aggregate bisection bandwidth of 51.2 Tb/s.
The centerpiece of this solution is the switch's integrated SHARPv3 (Scalable Hierarchical Aggregation and Reduction Protocol) technology. For the MoE workload, all-to-all communication was offloaded directly onto the switch fabric, with the switches performing in-network data reduction and redistribution. This eliminated the need for multiple host-side passes, dramatically reducing the communication overhead that had plagued the previous HDR deployment. The 920-9B210-00FN-0D0 datasheet highlights sub-100ns cut-through latency, which was validated during the proof-of-concept phase and proved critical for the MoE workload's stringent latency requirements.
The 920-9B210-00FN-0D0 specifications—including dual-redundant power supplies and hot-swappable fans—gave the team confidence in achieving their 99.999% availability target. The engineering team also confirmed that the 920-9B210-00FN-0D0 compatible ecosystem included all the OSFP optics and cables they had already qualified, eliminating procurement delays and simplifying the deployment timeline.
Results & Measurable Gains: From Bottleneck to Enabler
After the NDR fabric was fully deployed and the MoE training job restarted, the organization measured transformative improvements across multiple dimensions:
- GPU utilization recovery: Average GPU utilization surged from 52% to 89%, with peak utilization reaching 94% during compute-intensive phases—a direct result of eliminating network-induced stalls.
- All-to-all latency reduction: The time required for a full all-to-all exchange across 4,096 GPUs dropped from 180ms to under 45ms, a 75% improvement that directly translated to faster training iterations.
- Job completion time: The projected training timeline for the 1.8T MoE model was reduced from 14 days to just 9 days—a 36% acceleration that significantly cut both time-to-market and cloud compute costs.
- Operational efficiency: With 64 ports per switch, the number of spine switches required was reduced by 37% compared to a 40-port HDR design, simplifying cabling and reducing power consumption per rack by 28%.
From a total cost of ownership perspective, the organization's finance team noted that while 920-9B210-00FN-0D0 price per unit was higher than HDR alternatives, the per-GPU networking cost actually decreased due to the higher radix and reduced switch-layer count. This economic advantage, combined with the dramatic performance gains, made the NDR upgrade a clear business win. The 920-9B210-00FN-0D0 InfiniBand switch OPN solution also integrated seamlessly with their existing NVIDIA UFM deployment, providing centralized visibility and automated alerting that reduced operational overhead by 40%.
Summary & Outlook: The NDR Foundation for Next-Gen MoE Workloads
The NVIDIA Mellanox 920-9B210-00FN-0D0 proved to be the transformative solution this AI research organization needed to unlock the full potential of its 4,096-GPU H100 cluster. By delivering 400Gb/s NDR bandwidth, 64-port density, and SHARPv3 in-network computing, the switch enabled them to scale from 1,024 to 4,096 GPUs without compromising performance or manageability—a feat that would have been impossible with their previous HDR infrastructure.
Looking ahead, the organization plans to expand to 8,192 GPUs within the next 18 months, leveraging the 920-9B210-00FN-0D0 for sale through NVIDIA's channel partners to build an even larger three-tier folded-Clos fabric. For organizations evaluating their next-generation AI and HPC networking investments, the 920-9B210-00FN-0D0 offers a proven, production-ready solution that delivers on the promise of low-latency RDMA interconnect optimization at exascale.

