- By:
- Godoy, William F; Melnichenko, Tatiana A; Valero Lara, Pedro ; Elwasif, Wael R; Fackler, Philip W; Ferreira Da Silva, Rafael ; Teranishi, Keita ; Vetter, Jeffrey S
- Page Number:
- 2114-2128
- Book Title:
- Proceedings of the SC '25 Workshops of the International Conference for High Performance Computing, Networking, Storage and Analysis
- Publication Date:
- March 12, 2026
- Publisher Location:
- Association for Computing Machinery, New York, New York, United States of America
- Conference Name:
- The International Conference for High Performance Computing, Networking, Storage, and Analysis 2025 (SC '25)
- Conference Location:
- St. Louis, Missouri, United States of America
- Conference Sponsor:
- IEEE Computer Society and ACM SIGHPC and TCHPC
- View DOI Listing:
- https://doi.org/10.1145/3731599.3767573
Abstract
We explore the performance and portability of the novel Mojo language for scientific computing workloads on GPUs. As the first language based on the LLVM’s Multi-Level Intermediate Representation (MLIR) compiler infrastructure, Mojo aims to close performance and productivity gaps by combining Python’s interoperability and CUDA-like syntax for compile-time portable GPU programming. We target four scientific workloads: a seven-point stencil (memory-bound), BabelStream (memory-bound), miniBUDE (compute-bound), and Hartree–Fock (compute-bound with atomic operations); and compare their performance against vendor baselines on NVIDIA H100 and AMD MI300A GPUs. We show that Mojo’s performance is competitive with CUDA and HIP for memory-bound kernels, whereas gaps exist on AMD GPUs for atomic operations and for fast-math compute-bound kernels on both AMD and NVIDIA GPUs. Although the learning curve and programming requirements are still fairly low-level, Mojo can close significant gaps in the fragmented Python ecosystem in the convergence of scientific computing and AI.