- By:
- Tiwari, Devesh ; Gupta, Saurabh ; Gallarno, George; Rogers II, James H; Maxwell, Don E
- Page Number:
- 1
- Issue Number:
- N/A
- Book Title:
- Proceedings of the International Conference for High Performance Computing, Networking, Storage, and Analysis
- Publication Date:
- January 8, 2024
- Publisher Location:
- IEEE, New Jersey, United States of America
- Conference Name:
- Supercomputing (SC)
- Conference Location:
- Austin, Texas, United States of America
- View DOI Listing:
- https://doi.org/10.1145/2807591.2807666
Abstract
The high computational capability of graphics processing units (GPUs) is enabling and driving the scientific discovery process at large-scale. The world’s second fastest supercomputer for open science, Titan, has more than 18,000 GPUs that computational scientists use to perform scientific simu- lations and data analysis. Understanding of GPU reliability characteristics, however, is still in its nascent stage since GPUs have only recently been deployed at large-scale. This paper presents a detailed study of GPU errors and their impact on system operations and applications, describing experiences with the 18,688 GPUs on the Titan supercom- puter as well as lessons learned in the process of efficient operation of GPUs at scale. These experiences are helpful to HPC sites which already have large-scale GPU clusters or plan to deploy GPUs in the future.