A Learning-Based Reliability Platform for Automated GPU Fault Classification
Keywords:
.Abstract
Large-scale artificial intelligence (AI) training and inference systems rely on highly available acceleratorinfrastructure, where hardware faults directly reduce effective compute capacity and prolong recovery times. As GPU clusters grow in scale and architectural diversity
References
J. Dean, D. Patterson, and C. Young, “A New Golden Age in Computer Architecture,” Communications of the ACM, 2018. https://doi.org/10.1145/3282307
L. Barroso, U. Hölzle, and P. Ranganathan, The Datacenter as a Computer, 3rd ed., 2018.


