A Learning-Based Reliability Platform for Automated GPU Fault Classification

Authors

  • Kalyan Inturi

Keywords:

.

Abstract

Large-scale artificial intelligence (AI) training and inference systems rely on highly available acceleratorinfrastructure, where hardware faults directly reduce effective compute capacity and prolong recovery times. As GPU clusters grow in scale and architectural diversity

References

J. Dean, D. Patterson, and C. Young, “A New Golden Age in Computer Architecture,” Communications of the ACM, 2018. https://doi.org/10.1145/3282307

L. Barroso, U. Hölzle, and P. Ranganathan, The Datacenter as a Computer, 3rd ed., 2018.

https://doi.org/10.2200/S00874ED3V01Y201809CAC046

Downloads

Published

2026-01-15

How to Cite

Kalyan Inturi. (2026). A Learning-Based Reliability Platform for Automated GPU Fault Classification. Journal of Computational Analysis and Applications (JoCAAA), 35(1), 529–533. Retrieved from https://www.eudoxuspress.com/index.php/pub/article/view/4745

Issue

Section

Articles