Publications
Showing 30 results for Author: Jim H. Rogers II
Sep, 2026
ORNL Report
Challenges for Megawatt-Scale Artificial Intelligence Rack Infrastructure
As artificial intelligence (AI) computing densities continue to increase, industry is pursuing megawatt-scale rack architectures that require tightly coordinated advances in electrical power delivery, thermal management, operations, and infrastructure integration. This document summarizes the primary technical challenges for achieving this target.
Mar, 2026
Conference Paper
A Global Perspective on Supercomputer Power Provisioning: Case Studies from United States and Europe
Electrical provisioning in high performance computing is transitioning from simple nameplate Thermal Design Power (TDP) models to more nuanced approaches based on expected electrical load. This paper captures current power provisioning strategies across six international supercomputing centers and seven systems, three of which (Lumi, Summit, Sierra) were in the top 10 of t…
Jan, 2024
Conference Paper
Reliability Lessons Learned From GPU Experience With The Titan Supercomputer at Oak Ridge Leadership Computing Facility
The high computational capability of graphics processing units (GPUs) is enabling and driving the scientific discovery process at large-scale. The world’s second fastest supercomputer for open science, Titan, has more than 18,000 GPUs that computational scientists use to perform scientific simu- lations and data analysis. Understanding of GPU reliability characteristics, h…
Nov, 2023
Conference Paper
GPU Lifetimes on Titan Supercomputer: Survival Analysis and Reliability
The Cray XK7 Titan was the top supercomputer system in the world for a long time and remained critically important throughout its nearly seven year life. It was an interesting machine from a reliability viewpoint as most of its power came from 18,688 GPUs whose operation was forced to execute three rework cycles, two on the GPU mechanical assembly and one on the GPU circui…
Sep, 2019
Conference Paper
The Design, Deployment, and Evaluation of the CORAL Pre-Exascale Systems
CORAL, the Collaboration of Oak Ridge, Argonne and Livermore, is fielding two similar IBM systems, Summit and Sierra, with NVIDIA GPUs that will replace the existing Titan and Sequoia systems. Summit and Sierra are currently ranked No. 1 and No. 3, respectively on the Top500 list. We discuss the design and key differences of the systems. Our evaluation of the systems highl…
Dec, 2016
Journal
GPU Acceleration of the Locally Selfconsistent Multiple Scattering Code for First Principles Calculation of the Ground State and Statistical Physics of Materials
The Locally Self-consistent Multiple Scattering (LSMS) code solves the first principles Density Functional theory Kohn-Sham equation for a wide range of materials with a special focus on metals, alloys and metallic nano-structures. It has traditionally exhibited near perfect scalability on massively parallel high performance computer architectures. We present our efforts t…
Nov, 2016
Conference Paper
Using Balanced Data Placement to Address I/O Contention in Production Environments