Chips & SupplyUnited States
Upscale AI Partners With Nvidia While Competing With It

Every CPU maker always had its own and very proprietary way to link multiple CPUs into shared memory clusters and for the most part, this remains true despite the efforts of the CCIX consortium and even the CXL consortium to create a kind of universal glue to make many CPUs look like one much bigger CPU.
Attempts to create CPU SMP and NUMA standards largely failed because as much as customers would have liked to have a standard way to link CPUs together, the odds that they had a workload that would mix and match CPUs of different architectures from different vendors together seemed small. Moreover, each vendor could innovate on their SMP and NUMA chipsets, which were eventually embedded in the CPUs themselves, at a pace that was distinct from advances in their core designs, memory subsystems, and I/O. And most importantly, they were not dependent on some standards body to approve features they needed and they could differentiate as they saw fit.
So why do we think there will be a standard version of what amounts to a NUMA interconnect – what is commonly called scale up networking these days – for GPUs and XPUs? Because hope springs eternal, and because markets love competition because it drives down costs. (Substitution is a law of economics, not just a good idea....)
What is clear is that Nvidia can – and does – charge a hefty premium for its NVSwitch and NVLink ports, which offer unparalleled bandwidth and decent latency to glue as many as 72 GPUs to each other as a shared memory system that the company calls the NVL72 rackscale machine. With its NVLink Fusion effort, Nvidia is allowing those making custom CPUs to add NVLink ports to their devices so they can hook into the same memory space as Nvidia GPUs, and is similarly allowing XPU accelerators made by hyperscalers, cloud builders, and AI model builders to have NVLink ports so they can be put inside Nvidia’s rackscale architecture. Thus far, as far as we know, Nvidia is not allowing customers to add NVLink ports to both their own CPUs and their own XPUs and just buy NVSwitches and a license to the NVLink protocol from Nvidia to make shared memory systems.
That day might come, though, as I talked about in August in In The Long Run, Nvidia NVSwitch Is The InfiniBand Of Scale Up AI Networks. And in the longest of runs, should the right technology come to the market, Nvidia could adopt an industry standard scale up network with its own protocol overlays – particularly when it comes to AI inference. (There is no reason for Nvidia to yield such ground for AI training, where it has the vast majority of the market locked up.)
This is why Upscale AI, which I talked to back in January as it was raising $200 million in Series A funding, is so interesting. (I am not going to go over the history of the company here, and the serious players involved, so read that story if you want the background.) Upscale AI has raised $500 million in total so far, and its valuation has doubled to $2 billion and its employee count has also more than doubled to over 300 people since that January announcement. This week, the top brass at Upscale AI raised the curtain a little bit more on its plans for scale up and scale out networking aimed specifically at AI clusters.
Before you get too excited, we did not get much in the way of speeds and feeds for the “SkyHammer” scale up switch ASIC that has been in development for several years. Rajiv Khemani, co-founder and executive chairman at Upscale AI, and Barun Kar, its chief executive officer, have long and deep experience in high performance systems and networking, so we expect some interesting twists.
What we do know from reading that slide from a webinar – that’s Arvind Srikumar, senior vice president of product and marketing at the company, giving that part of the presentation – is that SkyHammer will have an aggregate bandwidth of 115.2 Tb/sec, which is state of the art and stands toe to toe with the best from Cisco Systems, Broadcom, and Nvidia for their scale out network ASICs. The chart also says that the SkyHammer chips and the SkyFabriX architecture it supports will eventually have switch gear that spans multiple petabits per second of aggregate bandwidth.
This will very likely be accomplished by ganging up multiple SkyHammer ASICs inside of a switch, much as Broadcom and Cisco often do for both fixed port and modular switches. As you know from watching the Broadcom and Meta Platforms presentations over the past decade in The Next Platform, you can double the bandwidth of ports with any given switch ASIC architecture by creating a baby two-tier network from six independent ASICs. Half the bandwidth of four of the ASICs are used to create downlinks, and the other half are used to cross-couple the four ASICs together using a pair of the same chips. This obviously doubles the latency for many of the port hops through the switch, but not for all of them. You can also use specialized fabric connectors (like the “Ramon” devices that are paired with Broadcom’s “Jericho” deep buffer ASICs) to create even larger modular switches.
Upscale AI will no doubt also create devices with higher aggregate bandwidth, and that is because the economics works out. You can build a switch with 2X the native bandwidth per port out of six switch ASICs, but it costs around 1.5X per port to do it compared to making a single device with 2X the bandwidth. (That is street pricing, not the cost of making the ASICs.) The company did not provide a roadmap for future SkyHammer ASICs.
As far as memory domain scalability goes, the SkyHammer chip’s architecture can support up to 576 accelerators in a single networking tier, and just like NVSwitch, using multiple network tiers it can scale further – in this case, to “thousands of accelerators.” How many thousands remains to be seen, but 1,024 or 1,156 seem to be the two numbers to aim for. The UALink consortium is promising 1,024 XPUs in a single tier, and Nvidia is going to try to scale NVSwitch beyond its current 576 GPU limit to 1,152 GPUs in the “Rubin Ultra” GPU generation. (This is with a multi-tier network, not a single tier.) Nvidia supports 72 GPUs in a single domain now, with 576 GPUs being possible with multiple tiers but not intended for production workloads.
If you do the math on that 576 devices number for SkyHammer, each device is getting 200 Gb/sec of bandwidth. This is obviously a lot less than the 1.8 TB/sec that an NVSwitch 5/NVLink 5 generation offers on the “Grace-Blackwell” NVL72 systems. The “Vera-Rubin” NVL72 systems will have NVSwitch 6/NVLink 6 with 3.6 TB/sec. With 72 XPUs, SkyHammer is delivering 1.6 Tb/sec per port, which is still only 200 GB/sec of memory bandwidth over the scale up network. That is one-sixth that of the NVSwitch 5 stack. (I have normalized the NVSwitch names
It is not at all clear that bandwidth matters more than latency, particularly for AI inference workloads. And predictable latency is always more important than low latency, as Google has taught us.
“One of the most important metrics when it comes to tokens is latency,” Srikumar tells The Next Platform. “We have to produce, but this is where it's a little bit different from what Rajiv and Barun has done in the past. With AI, the latency has to be super predictable, and we have built into the ASIC predictable jitter. And at the same time, transmission reliability within the domain is also important. We have some protocols which are standards-based that provide link level reliability, but within the fabric we have a mechanism to ensure the fabric is 100 percent lossless. This is why we call it reliable and predictable latency, where we are minimizing the latency. And you might be surprised with the numbers once we announce the specifics, If you look at a lot of Ethernet fabrics, there are pretty much a fixed pipeline arrangement, and it's a pipeline arrangement with less parallelism. Here we are doing a lot of things that are pushing the envelope on what has been done in the Ethernet fabrics to date. We believe SkyFabriX is going to be uniquely positioned in the market because we started with a clean sheet of paper rather than taking an Ethernet fabric and retrofitting the protocols that are needed for scale up.”
No matter what anyone believes, we are about to find out what works and what does not as people begin putting the UALink over Ethernet (UALoE in the common parlance today) and the ESUN memory protocols on top of fast Ethernet switches that have been tweaked to do scale up.
It is interesting to note that SkyHammer is running UALoE, not UALink proper, and the fact that it is running ESUN shows you that under the covers SkyHammer is an Ethernet switch ASIC. I was hoping that it was going to be a faster switch ASIC with more radix that could run UALink native and run an emulated version of ESUN when necessary. But, so it goes.
Here is the neat bit. Upscale AI has partnered with Nvidia to be its supplier of choice for scale out fabrics, and is choosing the Spectrum-X Ethernet switches in particular. Obviously not the Quantum-X InfiniBand switches, since the company’s founders are big believers in Ethernet eventually conquering InfiniBand. Some would say with the first generation of Ethernet switches that are compatible with the Ultra Ethernet Consortium’s specifications, this goal will be accomplished after nearly three decades of InfiniBand setting the high performance standard. InfiniBand will still beat the crap out of Ethernet when it comes to port-to-port hop latency. But Ethernet can get closer, and can scale out a lot further than InfiniBand for a given network topology. And Nvidia will be selling InfiniBand like crazy for those who need the lowest latency for their HPC and AI workloads. Fear not.
To be specific, Upscale AI will be acquiring Nvidia’s Spectrum-X ASICs and building its own scale-out switches from them as well as porting its SkyOS network operating system that runs on the scale up SkyFabriX switches so it also runs on the Spectrum-X ASICs. The reason this is possible is that SkyOS is an AI-optimized version of the SONiC network operating system and its related Switch Abstraction Interface (SAI) created by Microsoft way back in 2016 and which has become a de facto standard in many hyperscaler and cloud builder datacenters. The Upscale AI versions of Spectrum-X will support 400 Gb/sec, 800 Gb/sec, and 1.6 Tb/sec ports.
The important thing, according to Upscale AI, is that SkyOS will span both scale up and scale out networks, and its SkyCMD telemetry and orchestration tool will see both networks as a whole. The other important thing is that customers will be able to use any CPU or any GPU or any XPU front-ended by a DPU or NIC and will be able to mix and match these devices to create AI clusters. That is all well and good except that Nvidia is not supporting SkyHammer and its SkyOS as an alternative to NVSwitch for scale up. There is no way – as yet – to do scale up with Nvidia GPUs. AMD is actually counting on companies like Upscale AI to provide scale up networking options supporting UALink, UALoE, or ESUN, and it may lose patience and just acquire Upscale AI to own one of the options and make money from it. ( Why not? AMD is rich in stock funny money, and it may as well spend some of it.)
One last thing: Nvidia can always cut back on the bandwidth per port in an NVSwitch ASIC and increase the radix of the switch so it can flatten its NVSwitch network. But it will have to sacrifice bandwidth to do that. It might be faster to port the NVLink protocol to SkyHammer.... and have Nvidia become the dominant supplier of UALoE or ESUN switchery. We shall see.