- By:
- White, James
- Page Number:
- 461-467
- Book Title:
- SC Workshops '25: Proceedings of the SC '25 Workshops of the International Conference for High Performance Computing, Networking, Storage and Analysis
- Publication Date:
- March 12, 2026
- Publisher Location:
- Association for Computing Machinery (ACM), New York, New York, United States of America
- Conference Name:
- International Conference for High Performance Computing, Networking, Storage, and Analysis (SC '25), Workshop on Extreme Scale MPI
- Conference Location:
- St. Louis, Missouri, United States of America
- Conference Sponsor:
- Association for Computing Machinery
- View DOI Listing:
- https://doi.org/10.1145/3731599.3767389
Abstract
Near the full scale of exascale supercomputers, latency can dominate the cost of all-to-all communication even for very large message sizes. We describe GPU-aware all-to-all implementations designed to reduce latency for large message sizes at extreme scales, and we present their performance using 65536 tasks (8192 nodes) on the Frontier supercomputer at the Oak Ridge Leadership Computing Facility. Two implementations perform best for different ranges of message size, and all outperform the vendor-provided MPI_Alltoall. Our results show promising options for improving implementations of MPI_Alltoall_init.