March 2026

Conference Paper

A Performance-Portable MultiGPU Implementation of 3D Euler Equations using ProtoX and IRIS

By:
Mankad, Het Y; Monil, Mohammad Alaul Haque ; Rao, Sanil; Colella, Philip; Van Straalen, Brian; Franchetti, Franz; Vetter, Jeffrey S
Page Number:
1723-1731
Book Title:
SC24-W: Workshops of the International Conference for High Performance Computing, Networking, Storage and Analysis
Publication Date:
March 12, 2026
Publisher Location:
IEEE, New Jersey, United States of America
Conference Name:
15th Workshop on Latest Advances in Scalable Algorithms for Large-Scale Heterogeneous Systems at Super Computing Conference 2024
Conference Location:
Atlanta, Georgia, United States of America
Conference Sponsor:
CSMD
View DOI Listing:
https://doi.org/10.1109/SCW63240.2024.00215

Abstract

Computational scientists often face challenges when developing and optimizing code for high-performance computing (HPC), especially when trying to leverage GPUs. Given the heterogeneity of the nodes that comprise many modern HPC facilities, considerable demand exists for performance portable solutions for the core computational kernels used in many scientific computing libraries. In this work, we demonstrate a fourth-order finite volume method–based implementation of the Euler equations, which are an integral part of computational fluid dynamics. Our performance-portable multiGPU implementation for Euler equations uses ProtoX to generate kernels and IRIS for portability. ProtoX is a domain-specific language that uses a structured-grid partial differential equation library called Proto as its front end and the SPIRAL code generation system as its back end to generate optimized kernels for different architectures. Optimized kernels generated by ProtoX are orchestrated through the IRIS intelligent runtime system to provide portability. Two levels of optimizations within the IRIS runtime— directed acyclic graph fusion and task fusion—are explored to efficiently utilize computing resources in a multiGPU environment. Performance improvement through these optimizations is showcased by comparing the base ProtoX-IRIS implementation on AMD GPUs (Frontier node) and on NVIDIA GPUs (NVIDIA DGX-1).