The Three Puzzles of AI Data Center Fabric (Scale-Up, Out, Across)
Previous PostIn this, we shared the story of the birth of UET (Ultra Ethernet Transport), which emerged as an alternative to the existing slow TCP and cumbersome RoCEv2 to handle the explosive traffic of the artificial intelligence (AI) era.
However, many people often ask, “The concept is great, but is this really a feasible technology? When will it be installed in our data centers?” According to the latest report from global consulting group McKinsey & Company in February 2026, “By 2030, AI inference workloads will account for over 401 TP3T of data center demand and grow explosively by 351 TP3T annually.”It is called.
This means that next-generation AI networking is no longer a distant future imagination, but, The hottest reality and survival strategy already unfolding in data centers worldwide as of 2026It means...
Today, the three pillars of AI networking Scale-Up, Scale-Out, Scale-AcrossThrough these three puzzles, let's examine how HPE Networking's latest technology makes unraveling this complex infrastructure easy and powerful.
💡 A quick reminder! What is the difference between RoCEv2 and the next-generation AI Ethernet architecture?
Before we start putting the puzzle together in earnest, let's just point out one crucial difference between RoCEv2, which we covered last time, and this next-generation architecture.
Used in existing data centers as a cost-effective alternative to InfiniBand RoCEv2 (RDMA over Converged Ethernet)It was a studious style that was thoroughly obsessed with 'Lossless (zero packet loss)'.
Shall we take a look at the architecture structure below?
As you can see in the figure, in a RoCEv2 environment, when the road starts to get congested, a very busy congestion control loop (DCQCN) runs, such as the switch sending a Congestion Alert (ECN) to the destination and the destination sending a CNP packet to the source telling it to reduce the reception speed.
It even sends a 'Pause (PFC stop signal)' to the rear switch just before the buffer fills up completely, blocking the lane..
The problem is that the time it takes for this feedback to complete a full cycle and respond (RTT) is slower than the explosive speed of AI traffic. Consequently, this has chronically caused packet loss and 'PFC deadlocks,' where trailing traffic gets stuck one after another, paralyzing the entire road like a parking lot.
On the other hand, what we will talk about today Next-generation AI Ethernet architecture based on UETIt does not use the foolish method of “blocking the road to prevent loss.” Instead, it adopts a much more intelligent approach: “If the road is blocked, it splits the packets into smaller pieces and reroutes them in real-time to all available lanes (MRC/spray), and if it breaks, it recovers immediately at that section (LLR) rather than end-to-end.”.
Keeping this difference in mind, I will easily explain the three puzzles of AI networking using analogies.
🧩 Puzzle 1. Scale-Out: Solving 'Persistent Traffic Congestion' on a Road Filled with Dump Trucks
Where does it connect? Connections between servers and servers, and between switches.
AI training networks are characterized by tens of thousands of GPUs simultaneously exchanging a single massive chunk of data. In the networking world, this is called 'Elephant Flows,' and to use a simple analogy... A situation where the road is filled with magnificent large dump trucksno see.
The conventional navigation (Static ECMP routing) of existing data centers foolishly drove all these dump trucks into a single lane, often causing chronic congestion and packet loss.
The approach that big tech companies, including OpenAI, have recently focused on to solve this problem is precisely this. Packet Spray, that is, based on the UET standard Multi-path Reliable Connection (MRC) It is technology.
It is a method of breaking cargo (data) into very small packets and scattering them simultaneously across all empty lanes.
However, with this method, the delivery box depends on road conditions Number 1, number 5, number 3... All mixed up in formOut-of-OrderThere is a fatal side effect where ) arrives. Ultimately, to correct this distorted order, you are forced to use an expensive, dedicated, high-end NIC with 'hardware packet reordering capabilities' on the server side, which results in dependency on a specific vendor's proprietary technology and costs exploding.
In that case, is there no way to solve this chronic congestion using general-purpose infrastructure without proprietary technology? If you look at the evolution map of AI backend load balancing technology, you will see the answer.
AI Load Balancing Method
As you can see in the table, even NVIDIA is pushing the Ethernet-based 'Spectrum-X' packet spray method as its main focus to overcome the supply chain limitations and closed nature of InfiniBand, but there is still a barrier called Advanced NIC.
Here The unrivaled solution presented by HPE is RLB (RDMA-aware Load Balancing), or the 'Pinning' method, located in the third column of the graph.Instead of blindly splitting packets and spraying them, Juniper's intelligent algorithm accurately recognizes the flow of RDMA communication and A method of pre-designating (Pinning) the optimal lane with absolutely no congestion from the start, grouping them together, and sending them.no see.
Seeing is believing. Let's take a look at the actual benchmark performance data released at HPE Discover 2026.
This is the result of a test in which an extreme overload of 125% was intentionally applied to the network.
The existing navigation method (Static ECMP) experiences as many as 6,980 network congestion messages (ECN), causing the average bandwidth to drop sharply to 301 GB/s. On the other hand, the packet spray (DLB Per-Packet/MRC) method, which was theoretically perfect, only reached 358 GB/s due to the overhead of splitting and processing packets.
However, the one on the far right RLB Look at the environment. Overwhelming maximum speed with an average bandwidth of 366GB/sIt emits.
The most chilling part is the table at the bottom. Both the Congestion Message (ECN) and the Congestion Notification (CNP) Rate are '0'. It perfectly controlled network congestion and packet loss to 'zero' solely through the cleverness of the switch architecture, without forcing expensive dedicated NICs for packet reordering on the server side (Any NICs).
🧩 Puzzle 2. Scale-Up: AMD Helios AI Rack Breaks the 1 Trillion Parameter Barrier with Ethernet
Where does it connect? Connections between GPUs within a single server
Until now, GPUs have communicated at ultra-high speeds within a single server or rack. Scale-Up areaIt was an expensive and closed 'private exclusive road' like NVIDIA NVLink.. We couldn't install equipment from other companies, and the cost was incredibly high..
However, looking at the revenue forecasts for the Scale-Up Networking market below, you can see that the market landscape is rapidly shifting.
Source: 650 Group
If you look at the graph, just up until 2025, NVLink (proprietary technology), marked in orange, was practically monopolizing this market.. but Starting from the current year of 2026, the blue 'Open Ethernet' fabric sector is making explosive leaps, causing massive tectonic shifts.This is clearly visible..
And at the center of the fiercely rising blue graph is none other than HPE's AMD “Helios” AI Rack There is a solution.
Introduced through the collaboration of HPE and AMD AMD “Helios” AI RackIt is the reality of.
Instead of expensive and closed proprietary technology (NVLink), this architecture directly integrates the industry's first liquid-cooled Ethernet scale-up switch tray, shown in the top right, into the Compute system.
OCP (Open Compute Project) Standards and UEC (Ultra Ethernet Consortium) Specifications, and UALoE (Ultra Accelerator Link over Ethernet) By concentrating technology, we bundled as many as 72 large GPUs into a single open fabric and A phenomenal scale-up bandwidth of 260 TB/sWe achieved this with Ethernet..
If you read between the lines of this topographic map, you can see a very interesting stage of evolution. To move toward ESUN, the ultimate future standard, there is a prerequisite that the next-generation engine, the 'Tomahawk Ultra (TH-Ultra)' chipset, is required.
In order to break down the monopoly barriers of specific vendors and transition to open standards before that future arrives, The practical arena currently serving as the most powerful bridge is the Broadcom Tomahawk 6 (TH6)-based UALoE architecture.no see.
Industry First: Standard-based HPE Networking Ethernet Scale-up Switch
As you can see in the diagram above, inside this switch tray, cooling two TH6 chipsets 100% Direct-to-Chip Water Cooling TechnologyIt is equipped with safe PG25 refrigerant, as well as a backplane cartridge that bundles thousands of 200G cables without a single error.
Here HPE Juniper SONiC OS and AI-based intelligent NetOps softwareIt was combined to achieve extreme stability and observability.
A declaration to maximize AI computing efficiency on an open Ethernet highway where anyone can participate, instead of a specific vendor's closed membership road (NVLink).! This is the second one we got right Scale-UpIt is a puzzle..
🧩 Puzzle 3. Scale-Across: Breaking Down Borders and Data Centers with PTX·MX Routers
Where does it connect? Wide Area Network (WAN) interconnection between geographically distributed data centers, clouds, and users
The third and final puzzle goes beyond a single data center, connecting remote data centers or inference edge infrastructure located hundreds to thousands of kilometers away into one. Scale-Across It is an area.
According to research by global market research firm Omdia and HPE, AI training and inference traffic is at an annual average Explosive growth at an insane speed of over 140% (CAGR 140%+)It is continuing.
It is not simply that traffic is increasing. AI has now gone beyond the problems that can be solved within a single cluster. This is because there are four major realistic barriers that force companies to distribute their AI data centers across wide area networks (WANs).
Power and space limitations: The power consumption and rack space that a single data center can handle have reached their limits.
Computation and Memory Limits: It has become difficult to process a trillion-parameter model in a single place.
Latency and Reliability: We must provide ultra-low latency inference services in the user's vicinity.
Data Sovereignty and Regulation: To comply with national data regulations, data must be stored locally.
Ultimately, the following proposition holds true.
“The moment AI is geographically dispersed, the performance of the entire system is Routing and WAN networks determine the outcome."I do"” Once AI is distributed, performance becomes network-centric
If you look at the architecture map above, Scale-AcrossThe true nature of is clearly visible.
From large-scale distributed training between central AI data centers (Training over multi-DC) to the real-time deployment of updated models from the central data center to regional inference data centers (AI Inference DC) and carrier (SP) edge networks, all of these processes pass through the Wide Area Network (WAN).
However, standard Ethernet switches or InfiniBand cannot handle the sudden surge in AI traffic (Incast/Burst) the moment it travels over long distances, resulting in packet loss and delays.
The main players guarding the gateway to DC at this boundary and paving the way for long-distance highways are none other than MX and PTX Series Core Routersno see.
Inference Edge and External Network Gateway (MX Router): MX series (such as MX301) are placed at the forefront of Peering Edge and Inference DC gateways to securely and flexibly control traffic going out to enterprise and carrier networks.
Ultra-high-speed DCI and WAN backbone (PTX router): Delivering a whopping 36 ports of 800GE bandwidth in a 2RU form factor for Data Center Interconnection (DCI) sections PTX10002-36QD Place routers, etc. Large capacity Deep Buffer와 SRv6-based Traffic EngineeringWith this, it perfectly handles traffic bursts that occur during wide-area network communication.
Transponder-free Cost Innovation (Coherent Optics): By integrating next-generation coherent optical modules (JCO800 ZR/ZR+) that can transmit at 800G speeds up to 2,000km by simply plugging the module directly into the PTX router port without expensive and heavy optical transmission equipment (WDM), you can drastically reduce the cost of building DCI.
A technology that seamlessly connects the internal Open Ethernet (QFX) fabric to the Wide Area Core Routing Network (PTX/MX), eliminating physical barriers between data centers and scaling the AI computing fabric to a global scale! This is the third puzzle we have completed., Scale-Acrossno see.
Key Summary of Next-Generation AI Networking at a Glance
Today's topic The 3 Major Puzzles of AI Data Center FabricWe will organize it neatly into a table so that you can remember it at a glance.
division
Connection Area (Where)
Key Technology & Features (How)
everyday metaphors
HPE Juniper Core Solutions
Scale-Out
Server ↔ Switch (Backend Network)
UET/RoCEv2-based congestion control and packet retransmission minimization
State-of-the-art navigation that distributes deliveries in real-time while avoiding congested lanes
QFX5240 (800G) and unrivaled RLB algorithm
Scale-Up
GPU ↔ GPU (Inside rack/server)
UALoE-based Open Ethernet Scale-up and Direct-to-Chip Liquid Cooling Technology
A perfectly open highway with expensive monopoly tolls removed
AMD Helios AI Rack & QFX5250 (1.6T liquid-cooled switch tray)
Scale-Across
DC ↔ DC / WAN (Expanding to Earth Scale)
Large-capacity Deep Buffer, Traffic Engineering, and Coherent Optical Communication
High-capacity wide-area tunnels that directly connect continents without transponders
Ethernet is always “It is slow and vulnerable to packet loss.”It has been fighting against misunderstandings.
but Scale-OutIntelligent congestion control (RLB) in, Scale-UpOpen Ethernet rack infrastructure (AMD Helios & QFX5250) breaking the proprietary ecosystem (NVLink), and Scale-AcrossFrom to ultra-fast core routing (PTX/MX) that erases data center boundaries!
Through the open standard fabric completed by HPE, Ethernet is now a true, undisputed The most powerful protagonist of the AI data center backendIt has established itself as.
Will you be trapped on a closed, exclusive membership road, Or will it soar freely on the sustainable open Ethernet highway?
What does the architecture of the next-generation AI data center you are envisioning look like?