학술논문

Resiliency of HPC Interconnects: A Case Study of Interconnect Failures and Recovery in Blue Waters
Document Type
Periodical
Source
IEEE Transactions on Dependable and Secure Computing IEEE Trans. Dependable and Secure Comput. Dependable and Secure Computing, IEEE Transactions on. 15(6):915-930 Jan, 2018
Subject
Computing and Processing
Data security
Network security
Fault tolerance
Fault diagnosis
Multiprocessor interconnection
Data analysis
Networks
reliability
fault tolerance
fault diagnosis
Language
ISSN
1545-5971
1941-0018
2160-9209
Abstract
Availability of the interconnection network in high-performance computing (HPC) systems is fundamental to sustaining the continuous execution of applications at scale. When failures occur, interconnect recovery mechanisms orchestrate complex operations to recover network connectivity between the nodes. As the scale and design complexity of HPC systems increase, so does the system's susceptibility to failures during execution of interconnect-recovery procedures. This study characterizes the recovery procedures of the Gemini interconnect network, the largest Gemini network built by Cray, on Blue Waters, a 13.3 petaflop supercomputer at the National Center for Supercomputing Applications (NCSA). We propose a propagation model that captures interconnect failures and recovery procedures to help understand types of failures and their propagation in both the system and applications during recovery. The measurements show that recovery procedures occur very frequently and that the unsuccessful execution of recovery procedures, when additional failures occur during recovery, causes system-wide outages (SWOs, 28 out of 101) and application failures (3.4 percent of all running applications).