TechnologyUnited States

Arista steps up its Ethernet challenge to Nvidia’s AI scale-up networking fabric

Credited to Network World · networkworld.com

Useful
Photograph · Network World

Arista sees Ethernet as the scale-up networking fabric of choice for AI and continues to build up its alternative to proprietary interconnects such as Nvidia’s NVLink technology. This week, Arista rolled out new Ethernet reference architectures as well as a raft of integration partnerships with industry players—including AMD, Arm, Broadcom, Meta, Microsoft, and Qualcomm—aimed at helping large enterprises and hyperscale customers make Ethernet-based scale-up and scale-out designs a reality.

Arista will offer its 7060EX7 Series switches, along with fiber patch panels, integrated liquid-cooling manifolds, drip trays with active leak detection, rack management controllers, and power shelves with backup battery units, and combine that with software and other technologies from partner vendors.

All are based on standard Ethernet technologies: “Scale-up based on ESUN Ethernet, and scale-out based on Ethernet with UEC features,” Hardev Singh, vice president and general manager of Arista’s cloud and AI group, told Network World.

“ Arista’s positioning contrasts with Nvidia proprietary NVLink and InfiniBand stacks,” Singh said.

The Ethernet for Scale-Up Networking (ESUN) initiative was formed by the nonprofit Open Compute Project and promises to advance the networking technology to handle scale-up connectivity across accelerated AI infrastructure. ESUN backers include AMD, Arista, ARM, Broadcom, Cisco, HPE Networking, Marvell, Meta, Microsoft, Nvidia, OpenAI and Oracle.

ESUN provides network operators flexibility across accelerators, switching silicon, optics, physical rack designs, and network software while retaining the high bandwidth, low latency, and reliable in-order delivery required for collective communication, according to Arista.

The Ultra Ethernet Consortium ( UEC ), meanwhile, includes AMD, Arista, Broadcom, Cisco, Eviden, HPE, Intel, Meta, and Microsoft. Its charter is to bring together industry leaders to build a complete Ethernet-based communication stack architecture for high-performance networking.

“Arista’s strategy is that, until recently, Nvidia was only game in town, and their proprietary scale-up was the primary technology being used to build scale-up networks,” Singh said. “As the scale-up TAM is increasing, and the non-Nvidia accelerator ecosystem is increasing, that opens up a brand new opportunity for Arista. And that’s the other differentiator here—that this is all Ethernet. Strictly Ethernet.”

“You just need to look at what’s going on in the AI networking arena,” Singh said, citing Microsoft Maia, Meta’s MTIA (Meta Training and Inference Accelerator), AMD’s Helios, and Google’s Tensor Processing Units (TPU). “These technologies are all opportunities for Ethernet players like Arista,” Singh said.

The designs that Arista and its partners will offer will give customers a path to integrate compute, networking, power, and liquid cooling across evolving high-density AI deployments, Singh said. The three primary designs of what Arista calls Etherlink SU-144, an Ethernet ESUN-based scale-up architecture, include:

“The cross-rack architecture is particularly noteworthy because it extends the scale-up domain beyond the physical limits of a single rack, supporting up to 1,024 accelerators while maintaining high-bandwidth, low-latency connectivity,” wrote Bob Laliberte, principal analyst and founder at Liberte Research Group, in a LinkedIn post about the news.

“Arista is also addressing a broader infrastructure shift. As AI clusters become denser and move toward liquid cooling, the rack is replacing the individual switch as the practical unit of deployment. Power distribution, thermal management, fiber density, leak detection and system validation must now be engineered as an integrated solution,” Laliberte wrote.

In terms of scale-out, Arista is defining UEC-based designs that include Multipath Reliable Connection (MRC) fabric resiliency, multi-plane routing, intelligent dynamic load balancing, and cluster load balancing to handle AI token generation efficiency and workloads.

“Customer designs vary by accelerator power, GPU count, bandwidth, thermal envelope, and rack configuration,” Singh said. “Arista capabilities in chassis engineering, thermal design, and signal integrity [are] positioned as differentiators.”

For example, Singh said one rack configuration features bandwidth that scales with next-generation interconnects. A 16-switch configuration approaches 1.6 petabytes per rack, and other designs with XPO, CPO, and NPO technologies optical could quadruple bandwidth, Singh said. “A Tomahawk 7, 200G-based generation could support 6.5 petabytes in one rack exceeding 100 kilowatts.”

“Compared to today’s OSFP-based switches, which deliver 1.6 Pb/s per rack, next-generation optics increases that to 6.5 Pb/s per rack. This massive densification translates to nearly a 50% reduction in the overall data center footprint, resulting in significant time and financial savings,” according to Singh.

“So, I want to make it clear, customers are not going to find a rack SKU on Arista’s price list. We are not selling integrated racks. We are selling the switches that go into those liquid-cooled switches that go into the rack,” Sing said. “We’re offering our software and technology for integration with custom joint development designs, and we’re partnering with these partners, VARs or system integrators that the customers can work with to deliver the fully integrated rack. In addition, we’re working with different optics vendors as well as partners like Foxconn, Quanta, and Hive,” Singh said.

The foundation of all these designs is Arista’s Network Diagnostics Infrastructure (NetDI) diagnostics and validation layer, which runs between Arista’s EOS operating systems or SONiC, FBOSS, or other NOS. Arista’s newer 1.6T AI switches, for example, integrate NetDI specifically to provide consistent diagnostics across different operating environments.

NetDI can track several switch diagnostics functions such as cable and optic health, device telemetry, troubleshooting and failure isolation, and other factors that help cloud companies manage these massive scale networks, Singh said.

“NetDI offers deep hardware-level validation, signal integrity analysis, secure boot attestation, and Single Event Upset (SEU) resiliency, supporting rich telemetry for the switch, physical optics, power shelves, and liquid cooling infrastructure,” Arista wrote in a blog about the news.

Original · Network World

FrontMethod