Understanding Modern Neural Speech Recognition with Transformer Architectures

Authors

  • Uma Shankar Koushik Kethamakka

Keywords:

Transformer Architectures, Self-Supervised Learning, End-to-End Speech Recognition, Attention Mechanisms, Conformer, Streaming ASR, Multilingual Acoustic Modeling, Model Compression

Abstract

Neural speech recognition has been transformed by new transformer architectures, self-supervisedlearning and end-to-end training models. This article reviews the current state of the art systems from a representation learning perspective

References

Alec Radford et al., "Robust Speech Recognition via Large-Scale Weak Supervision," arXiv>eess>arXiv:2212.04356, 2022. [Online]. Available: https://arxiv.org/abs/2212.04356

Downloads

Published

2026-03-11

How to Cite

Uma Shankar Koushik Kethamakka. (2026). Understanding Modern Neural Speech Recognition with Transformer Architectures . Journal of Computational Analysis and Applications (JoCAAA), 35(3), 192–199. Retrieved from https://www.eudoxuspress.com/index.php/pub/article/view/5104

Issue

Section

Articles