TechnologyUnited States
ALCF Expands Aurora’s DAOS Storage to 800 Servers
Among the world’s fastest storage systems, DAOS delivers high performance across thousands of Aurora nodes to support data-intensive workloads ranging from massive simulations to AI model training.
Oct. 6, 2026 — Running simulations and AI workloads on one of the world’s most powerful supercomputers requires storage that can keep pace. On the ALCF ’s Aurora exascale system, Distributed Asynchronous Object Store (DAOS) provides the high-performance storage needed to move and manage massive amounts of data generated by large-scale scientific applications.
Aurora, the ALCF’s Intel-HPE exascale supercomputer, fully integrates DAOS with its compute fabric to support I/O and checkpointing for data-intensive scientific applications. Credit: Argonne.
The ALCF recently completed a major expansion of DAOS, bringing it from 128 storage servers initially, to 370 during an intermediate phase, and now to 800 in its final configuration. Extensive testing has demonstrated that DAOS is stable, highly performant, and able to scale effectively across thousands of Aurora compute nodes.
Available to Aurora users, DAOS has already established itself among the world’s fastest storage systems. Aurora remains No. 1 on the ISC26 IO500 Production list, based on benchmark results submitted in 2023 using an earlier DAOS configuration that demonstrated bandwidth exceeding 20 terabytes per second in individual tests. The ALCF team recently benchmarked the expanded system using all 800 servers which has pushed performance even higher, reaching data transfer rates of up to 30 terabytes per second using 2,048, 4,096, and 8,192 Aurora compute nodes. The testing shows that performance scales up and does not collapse under extreme client counts.
“We’ve put a lot of work into scaling DAOS and making sure it can deliver the performance and reliability our users need,” said Kevin Harms, ALCF-4 Technical Director. “The latest results show that DAOS can sustain high performance as applications scale across thousands of Aurora nodes, giving researchers a powerful resource to support data-intensive science.”
Write bandwidth as measured from HACC as a function of data type and file size. Credit: Adrian Pope, Argonne.
DAOS has also demonstrated strong performance for AI workloads. In the 2025 MLPerf Storage benchmark, a 128-server subset of Aurora’s DAOS system reached nearly 1 terabyte per second of write throughput and 600 gigabytes per second of read throughput, enabling a simulated checkpoint of Meta’s Llama 3 405B model in under 10 seconds. The results highlighted DAOS’s ability to support the intense I/O demands of large-scale AI training.
DAOS is an open-source object store designed specifically for highly parallel computing environments. On Aurora, it is directly connected to the supercomputer’s high-speed network, allowing applications to access the storage system at very high rates as they scale to larger portions of the machine. Its distributed architecture spreads data and metadata across the storage servers, helping provide high bandwidth and low latency without relying on a centralized metadata server. It also supports familiar interfaces used by scientific applications, including POSIX and MPI-I/O.
Researchers are already demonstrating what those capabilities can enable. In a recent effort on Aurora, an Argonne-led team used the Hardware/Hybrid Accelerated Cosmology Code (HACC) to complete two of the largest cosmological simulations ever performed.
The scale of the calculations created immense data-writing demands. The two simulations wrote more than 85 petabytes of data through DAOS, mostly as restart checkpoints that allowed the researchers to resume the calculations in the event of an interruption rather than start over. About 8 petabytes were retained for scientific analysis.
“Each run evolved more than 13 trillion particles across 8,100 Aurora nodes,” said Nicholas Frontiere, an Argonne computational scientist and member of the HACC team who led the simulation campaign on Aurora. “At that scale, I/O is a major computational challenge in its own right, and DAOS gave us the performance we needed to carry out these large-scale simulations.”
Earlier HACC testing had already demonstrated DAOS write speeds exceeding 5 terabytes per second on a smaller configuration of the storage system. The recent campaign shows how that capability can translate into production science at Aurora’s scale.
With its expanded capacity, demonstrated performance, and growing track record across simulation and AI workloads, DAOS gives Aurora users a high-performance option for applications with intensive storage requirements. To learn more about accessing and using DAOS on Aurora, visit the ALCF user guide.