March 2026

Conference Paper

Large-Message All-to-All Communication at Frontier Scale

By:
White, James
Page Number:
461-467
Book Title:
SC Workshops '25: Proceedings of the SC '25 Workshops of the International Conference for High Performance Computing, Networking, Storage and Analysis
Publication Date:
March 12, 2026
Publisher Location:
Association for Computing Machinery (ACM), New York, New York, United States of America
Conference Name:
International Conference for High Performance Computing, Networking, Storage, and Analysis (SC '25), Workshop on Extreme Scale MPI
Conference Location:
St. Louis, Missouri, United States of America
Conference Sponsor:
Association for Computing Machinery
View DOI Listing:
https://doi.org/10.1145/3731599.3767389

Abstract

Near the full scale of exascale supercomputers, latency can dominate the cost of all-to-all communication even for very large message sizes. We describe GPU-aware all-to-all implementations designed to reduce latency for large message sizes at extreme scales, and we present their performance using 65536 tasks (8192 nodes) on the Frontier supercomputer at the Oak Ridge Leadership Computing Facility. Two implementations perform best for different ranges of message size, and all outperform the vendor-provided MPI_Alltoall. Our results show promising options for improving implementations of MPI_Alltoall_init.


Related Researchers