Understanding Modern Neural Speech Recognition with Transformer Architectures
Keywords:
Transformer Architectures, Self-Supervised Learning, End-to-End Speech Recognition, Attention Mechanisms, Conformer, Streaming ASR, Multilingual Acoustic Modeling, Model CompressionAbstract
Neural speech recognition has been transformed by new transformer architectures, self-supervisedlearning and end-to-end training models. This article reviews the current state of the art systems from a representation learning perspective
References
Alec Radford et al., "Robust Speech Recognition via Large-Scale Weak Supervision," arXiv>eess>arXiv:2212.04356, 2022. [Online]. Available: https://arxiv.org/abs/2212.04356
Downloads
Published
2026-03-11
How to Cite
Uma Shankar Koushik Kethamakka. (2026). Understanding Modern Neural Speech Recognition with Transformer Architectures . Journal of Computational Analysis and Applications (JoCAAA), 35(3), 192–199. Retrieved from https://www.eudoxuspress.com/index.php/pub/article/view/5104
Issue
Section
Articles


