Multimodal Chatbots: Integrating Voice, Text, and Emotion Recognition for Human-Like Interactions

Authors

  • Sambasiva Rao Akkisetti,Satya Karteek Gudipati,Naveen Anand Mishra,Akbar Mohammed

Keywords:

Multimodal chatbots, emotion recognition, voice recognition, natural language processing (NLP), human-computer interaction

Abstract

Multimodal chatbots take that a step further, speaking in not just words but pictures, speech and even emotionalcues, producing — to some extent — interactions that feel more human. These chatbots utilize speech recognition,natural language processing (NLP) and emotion detection algorithms to understand user input and createmeaningful interactions. Voice recognition technology allows chatbots to understand various vocal features, suchas tone, pitch, and how quickly speak, and emotion recognition technologies build upon this feature by decoding

References

Rania Abdelghani, Yen-Hsiang Wang, Xingdi Yuan, Tong Wang, Pauline Lucas, Hélène Sauzéon, and Pierre-Yves Oudeyer. 2023. GPT-3-Driven Pedagogical Agents to Train Children’s Curious Question-Asking Skills. International Journal of Artificial Intelligence in Education (jun 2023). https://doi.org/10.1007/s40593- 023-00340-7

Downloads

Published

2025-04-15

How to Cite

Sambasiva Rao Akkisetti,Satya Karteek Gudipati,Naveen Anand Mishra,Akbar Mohammed. (2025). Multimodal Chatbots: Integrating Voice, Text, and Emotion Recognition for Human-Like Interactions. Journal of Computational Analysis and Applications (JoCAAA), 34(4), 1344–1352. Retrieved from https://www.eudoxuspress.com/index.php/pub/article/view/3303

Issue

Section

Articles