March 2026

Conference Paper

Sequence length scaling in vision transformers for scientific images on frontier

By:
Tsaris, Aristeidis ; Wang, Feiyi ; Balaprakash, Prasanna ; Wang, Xiao ; Lu, Dan ; Yin, Junqi ; Ashfaq, Moetasim ; Liu, Siyan ; Choi, Jong Youl ; Fan, Ming ; Zhang, Chengming; Mohamed, Wahib
Journal Name:
The International Journal of High Performance Computing Applications
Volume:
TBD
Publication Date:
March 12, 2026
Conference Name:
The International Conference for High Performance Computing, Networking, Storage, and Analysis (SC'24)
Conference Location:
Atlanda, Georgia, United States of America
Conference Sponsor:
IEEE
View DOI Listing:
https://doi.org/10.1177/10943420251394758

Abstract

Vision Transformers (ViTs) are pivotal for foundational models in scientific imagery, including Earth science applications, due to their capability to process large sequence lengths. While transformers for text have inspired scaling sequence lengths in ViTs, adapting these for ViTs introduces unique challenges. We develop distributed sequence parallelism for ViTs, enabling them to handle up to 1M tokens. Our approach, leveraging DeepSpeed-Ulysses and Long-Sequence-Segmentation with model sharding, is the first to apply sequence parallelism in ViT training, achieving a 94% batch scaling efficiency on 2,048 AMD-MI250X GPUs. Evaluating sequence parallelism in ViTs, particularly in models up to 10B parameters, highlighted substantial bottlenecks. We countered these with hybrid sequence, pipeline, and flash attention strategies, to scale beyond single GPU memory limits. Our method significantly enhances climate modeling accuracy by 20% in temperature predictions, marking the first training of a vision transformer model to convergence with a sequence length of 188K tokens, using full self-attention.