- By:
- Engelmann, Christian ; Vallee, Geoffroy R; Naughton III, Thomas J; Scott, Stephen L
- Page Number:
- 252-257
- Book Title:
- Proceedings of the 17th Euromicro International Conference on Parallel, Distributed, and network-based Processing (PDP) 2009
- Publication Date:
- February 20, 2009
- Publisher Location:
- IEEE Computer Society, Los Alamitos, California, United States of America
- Conference Name:
- 17th Euromicro International Conference on Parallel, Distributed, and network-based Processing (PDP) 2009
- Conference Location:
- Weimar, Germany
- View DOI Listing:
- https://doi.org/10.1109/.30
Abstract
Proactive fault tolerance (FT) in high-performance computing is a concept that prevents compute node failures from impacting running parallel applications by preemptively migrating application parts away from nodes that are about to fail. This paper provides a foundation for proactive FT by defining its architecture and classifying implementation options. This paper further relates prior work to the presented architecture and classification, and discusses the challenges ahead for needed supporting technologies.