March 2026

Conference Paper

Characterizing the Impact of GPU Power Management on an Exascale System

By:
Costa, Mariana; Georgiadou, Antigoni ; White, James ; Alvarez, Bruno; Polo, Jorda; Shin, Woong ; Navaux, Philippe; Messer II, Otis E; Lorenzon, Arthur
Page Number:
1524-1533
Book Title:
Proceedings of the SC '25 Workshops of the International Conference for High Performance Computing, Networking, Storage and Analysis
Publication Date:
March 12, 2026
Conference Name:
2025 The International Conference for High Performance Computing, Networking, Storage, and Analysis (SC25), Sustainable Supercomputing Workshop
Conference Location:
St. Louis, Missouri, United States of America
Conference Sponsor:
ACM, SIGHPC, IEEE Computer Society, TCHPC
View DOI Listing:
https://doi.org/10.1145/3731599.3767702

Abstract

As GPU-accelerated high-performance computing (HPC) systems approach exascale performance, controlling energy consumption without compromising throughput is essential. Architectures such as the AMD MI250X-based Frontier supercomputer provide runtime mechanisms like frequency and power capping, enabling energy tuning without modifying application code. Although both target energy reduction, they operate via distinct hardware control paths and influence workloads differently. We present a comprehensive evaluation of these strategies on a leadership-class system using diverse HPC proxy applications representative of production workloads. Our study analyzes performance–energy trade-offs across multiple capping levels, node counts (1 and 32), and application profiles. Results show that frequency capping generally achieves higher energy efficiency and scalability, with gains of up to 13.2% without performance loss, while power capping is more effective for single-node runs or bursty GPU utilization. We also provide practical guidelines to help system administrators and users balance energy efficiency and performance in large-scale scientific workloads.