January 2024

Conference Paper

Reliability Lessons Learned From GPU Experience With The Titan Supercomputer at Oak Ridge Leadership Computing Facility

By:
Tiwari, Devesh ; Gupta, Saurabh ; Gallarno, George; Rogers II, James H; Maxwell, Don E
Page Number:
1
Issue Number:
N/A
Book Title:
Proceedings of the International Conference for High Performance Computing, Networking, Storage, and Analysis
Publication Date:
January 8, 2024
Publisher Location:
IEEE, New Jersey, United States of America
Conference Name:
Supercomputing (SC)
Conference Location:
Austin, Texas, United States of America
View DOI Listing:
https://doi.org/10.1145/2807591.2807666

Abstract

The high computational capability of graphics processing units (GPUs) is enabling and driving the scientific discovery process at large-scale. The world’s second fastest supercomputer for open science, Titan, has more than 18,000 GPUs that computational scientists use to perform scientific simu- lations and data analysis. Understanding of GPU reliability characteristics, however, is still in its nascent stage since GPUs have only recently been deployed at large-scale. This paper presents a detailed study of GPU errors and their impact on system operations and applications, describing experiences with the 18,688 GPUs on the Titan supercom- puter as well as lessons learned in the process of efficient operation of GPUs at scale. These experiences are helpful to HPC sites which already have large-scale GPU clusters or plan to deploy GPUs in the future.


Related Researchers