Academic Journal

A review on Gujarati language based automatic speech recognition (ASR) systems.

Λεπτομέρειες βιβλιογραφικής εγγραφής
Τίτλος: A review on Gujarati language based automatic speech recognition (ASR) systems.
Συγγραφείς: Dua, Mohit, Bhagat, Bhavesh, Dua, Shelza, Chakravarty, Nidhi
Πηγή: International Journal of Speech Technology; Mar2024, Vol. 27 Issue 1, p133-156, 24p
Θεματικοί όροι: Automatic speech recognition, Language models, Feature extraction, Extraction techniques, Speech
Περίληψη: Automatic speech recognition (ASR) plays a crucial role in facilitating natural and efficient human–computer interaction. This paper offers a comprehensive review of ASR systems tailored specifically for the Gujarati language. Existing literature on Gujarati ASR indicates that most survey papers have focused solely on back-end classification methods. This paper aims to fill this gap by presenting a comprehensive survey that encompasses feature extraction techniques, backend models, speech datasets, and evaluation metrics. This review provides an in-depth analysis of the cutting-edge methodologies employed in the development of Gujarati ASR systems. Firstly, the paper discusses various feature extraction techniques. Secondly, it covers classification models and their impact on ASR performance. Thirdly, the study delves into available speech datasets relevant to the Gujarati language, offering valuable insights for researchers and practitioners. Finally, the paper reviews the evaluation parameters used in ASR systems. Additionally, it gives an overview of online toolkits, resources, and language models pertinent to Gujarati ASR. This review serves as a comprehensive reference for academics and professionals involved in the research and development of ASR systems for the Gujarati language. The parameters selected for comparison include front-end feature extraction methods, back-end classification techniques, speech dataset, and evaluation metrics. Furthermore, the paper discusses both the contributions and limitations encountered by current ASR systems during the review. Finally, it addresses various challenges that still persist and provides directions for future research in this critical field. [ABSTRACT FROM AUTHOR]
Copyright of International Journal of Speech Technology is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Βάση Δεδομένων: Complementary Index
Περιγραφή
ISSN:13812416
DOI:10.1007/s10772-024-10087-8