Data Center Switches for Gen AI/ML

Artificial intelligence (AI) is being used across all industries and sectors, becoming a universal requirement. AI-based solutions are essential for analyzing continuously incoming data, gaining insights, and supporting real-time decision-making using task-specific AI models, such as machine learning and natural language processing (NLP).

인공지능 사용 사례

Infrastructure Requirements for GenAI (Generative AI)

Large-scale GenAI models require massive amounts of GPU-powered computing to support terabyte-sized training datasets containing billions of parameters accessed through ultra-low-latency memory and storage systems. The networks supporting these large-scale GenAI models are tailored to deliver optimized, power-efficient, and predictable performance across LLM multi-tenant workload environments.

The fabric and topology are highly tuned to support the entire system, with all aspects of the network including compute, data processing units (DPUs), I/O, cabling and optics, acceleration software, and networking (e.g., InfiniBand 400/800G Ethernet switches).

Preparing Data Center Networks for AI

AI architectures require a dedicated network fabric that combines high-performance and low-latency connectivity to ensure the fastest training, inference, and modeling task completion times. Early HPC and AI training networks initially required high-speed, low-latency, and dedicated connectivity for fast and efficient communication between servers and storage systems. InfiniBand The network gained popularity.

Today, 100/200/400G+ Leaf/Spine Ethernet switching is gaining significant traction in supporting networking for HPC/AI clusters by providing an open, standards-based alternative, and is expected to become the most cost-effective and popular alternative for many AI use cases.

AI 네트워킹 시장

“As AI bandwidth increases, the share of Ethernet switching connected to AI/ML and accelerated computing will shift to a significant portion of the market by 2027. As AI/ML-enabled products reach production scale, shipments of 800Gbps-based switches and optical modules will reach record levels.”

Alan Weckel, founder and technology analyst at 650 Group

Modern AI applications require high-bandwidth, lossless, low-latency, and scalable multi-tenant networks interconnecting hundreds or thousands of GPUs at speeds ranging from 100G to 400G and beyond. An Ethernet-based networking fabric delivers the reliability and performance required for AI workload clusters with hundreds or thousands of GPUs.

Backend vs. Frontend Differences in AI Fabric Networks

  • Frontend network: It is built using existing Ethernet networks and is built as a lossless environment to support shared storage.

  • GenAI Fabric: Workloads where the GPU moves workload traffic using both the internal PCIe bus and the backend fabric.

  • Backend network: Used to interconnect AI servers and workloads, and may also include dedicated storage. Built with no link oversubscription between workloads and the fabric, and optimized for GPU workloads using low latency, DCB, RoCE, etc.

GenAI Backend Fabric Network Design Method
Multi-layer Clos
  • ToR switches connect each server and provide connectivity to other racks through aggregation switches.
  • Spine switches provide connections to other PODs
  • Best suited for CPU-intensive workloads
Rail Optimized
  • Built on a GPU-centric cluster with two different communication paths to each GPU.
  • One path is NVIDIA® NVSwitch1Through, other routes connect via rail switches.
  • NVIDIA NVSwitch on individual servers creates high-speed interconnects to form high-bandwidth (HB) domains.
  • These rail switches are connected to the Spine switches to form a full-section, any-to-any Clos network topology.
Rail Only
  • Similar to Rail Optimized
  • Instead, remove network connections between GPUs of different ranks within the Rail.
  • HB domain2Communication is still possible by passing data through
Rail-only architecture

AI-enabled data center switches from HPE Aruba Networking

HPE Aruba Networking can help you design and build a dedicated AI network fabric.

HPE Aruba CX 8325 Switch Series

The CX 8325 switch series is an entry-level GenAI switch solution that supports 1/10/25/100GbE ports.
It supports rail-only architecture and is suitable for general data center environments as a 1U switch with up to 6.4Tbps.

HPE Aruba CX 9300 Switch Series

CX 9300 Switch Seriesis a next-generation 25.6Tbps, 1U switch that supports 32 100/200/400GbE ports.

HPE Aruba Networking CX9300-32D

The CX 9300 offers AI/HPC-optimized features, including low latency, lossless network quality of service (QoS), ROCEv2, ECN, and PFC, all of which are essential for AI/HPC connectivity. It also supports both rail-only and rail-optimized architectures, and the CX 9300 switch can be used for either Spine or Rail purposes.


For more information, please see the link below.


  1. Supports high bandwidth but short-distance interconnection ↩︎
  2. Subset of GPU ↩︎