March 2026

Conference Paper

Comparative Study of Large Language Model Architectures on Frontier

By:
Yin, Junqi ; Bose, Avishek ; Cong, Guojing ; Lyngaas, Isaac R; Anthony, Quentin
Page Number:
556-569
Book Title:
2024 IEEE International Parallel and Distributed Processing Symposium (IPDPS)
Publication Date:
March 12, 2026
Publisher Location:
IEEE, New Jersey, United States of America
Conference Name:
IEEE International Parallel and Distributed Processing Symposium (IPDPS)
Conference Location:
San Francisco, California, United States of America
Conference Sponsor:
IEEE
View DOI Listing:
https://doi.org/10.1109/IPDPS57955.2024.00056

Abstract

Large language models (LLMs) have garnered significant attention in both the AI community and beyond. Among these, the Generative Pre-trained Transformer (GPT) has emerged as the dominant architecture, spawning numerous variants. However, these variants have undergone pre-training under diverse conditions, including variations in input data, data preprocessing, and training methodologies, resulting in a lack of controlled comparative studies. Here we meticulously examine two prominent open-sourced GPT architectures, GPT-NeoX and LLaMA, leveraging the computational power of Frontier, the world’s first Exascale supercomputer. Employing the same materials science text corpus and a comprehensive end-to-end pipeline, we conduct a comparative analysis of their training and downstream performance. Our efforts culminate in achieving state-of-the-art performance on a challenging materials science benchmark. Furthermore, we investigate the computation and energy efficiency, and propose a computationally efficient method for architecture design. To our knowledge, these pre-trained models represent the largest available for materials science. Our findings provide practical guidance for building LLMs on HPC platforms.