NVIDIA Mellanox 920-9B210-00FN-0D0 in Practice: Building an NDR Low-Latency Fabric for Trillion-Parameter AI Clusters
August 26, 2026
NVIDIA Mellanox 920-9B210-00FN-0D0 in Practice: Building an NDR Low-Latency Fabric for Trillion-Parameter AI Clusters
Background & The Challenge: When HDR Meets Trillion-Parameter Models
A leading AI research organization recently scaled its large language model training cluster from 1,024 to 4,096 NVIDIA H100 GPUs, with the goal of compressing the training timeline for a 1.8-trillion-parameter sparse MoE (Mixture of Experts) model from months to under two weeks. However, during initial validation, the team observed severe congestion in the all-to-all communication phase on their existing 200Gb/s HDR fabric, causing GPU utilization to plummet from an expected 85% to just 52%—far below the threshold required to meet their ambitious training deadline.
The network analysis revealed that the all-to-all collective operations, which are intrinsic to MoE architectures, were saturating the HDR links and triggering congestion control mechanisms that further amplified latency. The organization needed a fabric that could deliver deterministic, sub-microsecond latency across thousands of endpoints while providing the bandwidth headroom to absorb traffic bursts without performance collapse. This is precisely the challenge that the NVIDIA Mellanox 920-9B210-00FN-0D0 was designed to address.
Solution & Deployment: A 400Gb/s NDR Fabric for Exascale AI
The organization deployed the 920-9B210-00FN-0D0 InfiniBand switch OPN (Ordering Part Number) as the foundation of a new two-tier leaf-spine fabric. At the leaf layer, each rack received two 920-9B210-00FN-0D0 MQM9790-NS2F 400Gb/s NDR switches, providing 128 ports of 400Gb/s connectivity to 64 dual-port H100 servers per rack via OSFP-to-OSFP active optical cables. At the spine layer, 16 additional 920-9B210-00FN-0D0 switches interconnected all leaf switches, creating a non-blocking fat-tree with full bisection bandwidth of 51.2 Tb/s per spine tier.
What made this solution particularly compelling was the switch's integrated SHARPv3 (Scalable Hierarchical Aggregation and Reduction Protocol) technology. For the MoE workload, all-to-all communication was offloaded directly onto the switch fabric, with the switches performing in-network data reduction and redistribution. This eliminated the need for multiple host-side passes, dramatically reducing the communication overhead that had plagued the previous HDR deployment.
According to the 920-9B210-00FN-0D0 datasheet, the switch delivers sub-100ns cut-through latency, which was validated during the proof-of-concept phase. The engineering team also confirmed that the 920-9B210-00FN-0D0 compatible ecosystem included all the OSFP optics and cables they had already qualified, eliminating procurement delays. The 920-9B210-00FN-0D0 specifications—including dual-redundant power supplies and hot-swappable fans—gave the team confidence in achieving their 99.999% availability target for the production cluster.
Results & Measurable Gains: From Bottleneck to Enabler
After the full NDR fabric was deployed and the MoE training job was restarted, the organization measured transformative improvements across multiple dimensions:
- GPU utilization recovery: Average GPU utilization surged from 52% to 89%, with peak utilization reaching 94% during compute-intensive phases—a direct result of eliminating network-induced stalls.
- All-to-all latency reduction: The time required for a full all-to-all exchange across 4,096 GPUs dropped from 180ms to under 45ms, a 75% improvement that directly translated to faster training iterations.
- Job completion time: The projected training timeline for the 1.8T MoE model was reduced from 14 days to just 9 days—a 36% acceleration that significantly cut both time-to-market and cloud compute costs.
- Operational efficiency: With 64 ports per switch, the number of spine switches required was reduced by 37% compared to a 40-port HDR design, simplifying cabling and reducing power consumption per rack by 28%.
From a total cost of ownership perspective, the organization's finance team noted that while the 920-9B210-00FN-0D0 price per unit was higher than the HDR alternative, the per-GPU networking cost actually decreased due to the higher radix and reduced switch-layer count. This economic advantage, combined with the dramatic performance gains, made the NDR upgrade a clear business win. The 920-9B210-00FN-0D0 InfiniBand switch OPN solution also integrated seamlessly with their existing NVIDIA UFM deployment, providing centralized visibility and automated alerting that reduced operational overhead by 40%.
Summary & Outlook: The NDR Foundation for Next-Gen AI
The NVIDIA Mellanox 920-9B210-00FN-0D0 proved to be the transformative solution this AI research organization needed to unlock the full potential of its 4,096-GPU H100 cluster. By delivering 400Gb/s NDR bandwidth, 64-port density, and SHARPv3 in-network computing, the switch enabled them to scale from 1,024 to 4,096 GPUs without compromising performance or manageability—a feat that would have been impossible with their previous HDR infrastructure.
Looking ahead, the organization plans to expand to 8,192 GPUs within the next 18 months, leveraging the 920-9B210-00FN-0D0 for sale through NVIDIA's channel partners to build an even larger three-tier folded-Clos fabric. For organizations evaluating their next-generation AI and HPC networking investments, the 920-9B210-00FN-0D0 offers a proven, production-ready solution that delivers on the promise of low-latency RDMA interconnect optimization at exascale.

