Publications

Showing 22 results for Author: Trey White

  • Sep, 2026

    Conference Paper

    Power management methods such as power capping or dynamic frequency capping have long been considered primary mechanisms for addressing the power constraints of exascale computing and beyond. However, their effectiveness in GPU-accelerated HPC applications is not straightforward, as they are inherently multi-kernel and exhibit heterogeneous computational and memory behavio…

  • Aug, 2026

    Conference Paper

    Power is a fundamental constraint as supercomputing advances to exascale. Efficient operation within strict power budgets requires application-aware power management based on a detailed understanding of application-level power behavior. This work analyzes the Energy Exascale Earth System Model (E3SM) atmosphere component, SCREAM, on Perlmutter (NERSC) and Frontier (OLCF).…

  • Mar, 2026

    Conference Paper

    As GPU-accelerated high-performance computing (HPC) systems approach exascale performance, controlling energy consumption without compromising throughput is essential. Architectures such as the AMD MI250X-based Frontier supercomputer provide runtime mechanisms like frequency and power capping, enabling energy tuning without modifying application code. Although both target…

  • Mar, 2026

    Conference Paper

    Dragonfly-based networks are an extensively deployed network topology in large-scale high-performance computing due to their cost-effectiveness and efficiency. The US will soon have three Exascale supercomputers for leadership class workloads deployed using dragonfly networks. Compared to indirect networks of similar scale, the dragonfly network has considerably reduced ca…

  • Mar, 2026

    Conference Paper

    Near the full scale of exascale supercomputers, latency can dominate the cost of all-to-all communication even for very large message sizes. We describe GPU-aware all-to-all implementations designed to reduce latency for large message sizes at extreme scales, and we present their performance using 65536 tasks (8192 nodes) on the Frontier supercomputer at the Oak Ridge Lead…

  • Mar, 2026

    Conference Paper

    As computing systems approach the limits of traditional silicon technology, the diminishing returns in performance per watt present a significant barrier to sustaining growth in HPC. From a large-scale scientific supercomputing facility point of view, we propose a multifaceted strategy toward specialized hardware and architectures that are optimized for energy efficiency i…

  • Mar, 2026

    Conference Paper

    Climate simulation will not grow to the ultrascale without new algorithms to overcome the scalability barriers blocking existing implementations. Until recently, climate simulations concentrated on the question of whether the climate is changing. The emphasis is now shifting to impact assessments, mitigation and adaptation strategies, and regional details. Such studies wil…

  • Mar, 2026

    Conference Paper

    The Simple Cloud-Resolving E3SM Atmosphere Model (Scream) won the inaugural ACM Gordon Bell Prize for Climate Modeling. While most of Scream is portable Kokkos code, the Gordon-Bell runs did include tuning specifically for Frontier, the exascale computer at Oak Ridge National Laboratory. Production science runs use the same high-level configuration of Scream, but the tuned…

  • Jan, 2026

    Conference Paper

    Advanced ab initio materials simulations face growing challenges as increasing systems and phenomena complexity requires higher accuracy, driving up computational demands. Quantum many-body GW methods are state-of-the-art for treating electronic excited states and couplings but often hindered due to the costly numerical complexity. Here, we present innovative implementatio…

1
2