Power & SitesUnited States

National Compute Grid Launches to Pool Idle AI Capacity Across Providers

Credited to HPCwire · hpcwire.com

Useful

For all the talk about a shortage of AI computing power, and the need for more powerful chips and bigger data centers, there’s another side to the story. A significant amount of existing capacity isn’t being used. While tech companies race to build massive new data centers, expensive GPUs in other facilities can sit idle between workloads.

What if all that fragmented computing capacity could be somehow pooled together to become something much bigger. That is the goal of an ambitious new initiative called the National Compute Grid. It is a nationwide network of companies that are capable of pooling GPU resources and making them available wherever they’re needed. At least that is the concept behind it.

The idea is that instead of always building more data centers or buying more powerful chips, why not make better use of the computing resources that are already out there?

Here’s how it would work. Data centers and cloud providers can make their unused computing capacity available through a software platform called Grid Exchange. The network is available to AI companies, startups, researchers, and other users who can find the GPUs they need. The Grid will match them to available computing resources. It takes into account the type of hardware required, available capacity and what customers are willing to pay.

(Shutterstock/PHOTOCREO-Michal-Bednarek)

The grid is designed to support different chip architectures, including GPUs from NVIDIA and AMD and TPUs from Google. That gives customers more options. However, moving workloads between different types of hardware isn’t always straightforward because of software compatibility issues.

The scale of the project is already substantial. At least on paper. According to its organizers, the grid has around 760 megawatts of computing capacity either connected or in the pipeline. The goal is to reach 2 gigawatts by 2030.

According to the organizers, single-tenant data centers use less than 15% of their computing capacity on average. That doesn’t mean 85% of all GPUs are sitting idle, but it does show how much computing power may be going unused at some facilities.

This could be particularly useful for smaller AI companies. Unlike OpenAI and Anthropic, which can afford to sign big, long-term computing contracts, smaller companies don’t always have that luxury. Sam Sinha, head of AI at robotics company 1X, spoke about this problem to Axios. For a smaller developer that needs extra GPUs for a few weeks or a particular project, having somewhere else to turn could be a big help.

Government agencies, academic institutions, and national labs can use it too. Some research projects require more computing power than an institution has available, and getting time on existing systems isn’t always easy. This may provide the much needed computing capacity required by researchers.

The grid could also be useful for President Trump’s Genesis Mission, which needs a lot of computing power for AI research. National Compute plans to contribute $100 million in computing credits to the mission, so researchers would have another way to get access to the GPUs they need.

Participating researchers can tap into computing resources available across the grid. This helps reduce how much the users depend on government supercomputers or dedicated cloud contracts. It would also give National Compute an opportunity to show that its model can work for scientific workloads – not just commercial AI applications.

The timing also matters. The announcement for the initiative comes as the White House prepares to announce more than $1 billion in private-sector commitments to the Genesis Mission. Companies including AMD, OpenAI and Anthropic are expected to take part.

The initiative is being spearheaded by Anjney Midha, who leads the AI holding company Amp. Cloud infrastructure provider Vultr is a confirmed founding member of the consortium and is helping bring computing capacity to the network. National Compute Labs is developing the grid and its exchange platform. Other AI companies, infrastructure operators and investors are involved, although a complete list of members has not been made public.

The consortium has taken a different approach to pricing. According to Vultr, the grid will support a model called “goodput,” where customers are charged for completed, uninterrupted computing work. This is different to just pricing it based on the number of GPU-hours used. National Compute has reportedly attracted more than $5 billion in reservations already. However, that should not be confused with revenue already earned.

There will be challenges ahead, and that is expected. After all, bringing together GPUs from different providers is one thing. Making sure they can reliably handle demanding AI workloads is another challenge entirely.

Training large AI models often requires thousands of GPUs connected through extremely fast networks. Having spare GPUs spread across multiple data centers doesn’t necessarily mean they can be used together for those jobs. Unfortunately, it’s not that simple. Moving large datasets between facilities can also be expensive and time-consuming.

The grid may be more useful initially for smaller training runs, inference and research workloads that don’t require thousands of GPUs working together. Whether it can reliably support larger jobs remains to be seen.

Original · HPCwire

FrontMethod