Papers › Is Attention always needed? A Case Study on Language Identification from Speech

Is Attention always needed? A Case Study on Language Identification from Speech

5 Oct 2021arXiv:2110.03427archive 2025-07-28

Atanu Mandal, Santanu Pal, Indranil Dutta, Mahidas Bhattacharya, Sudip Kumar Naskar

Language Identification (LID) is a crucial preliminary process in the field of Automatic Speech Recognition (ASR) that involves the identification of a spoken language from audio samples. Contemporary systems that can process speech in multiple languages require users to expressly designate one or more languages prior to utilization. The LID task assumes a significant role in scenarios where ASR systems are unable to comprehend the spoken language in multilingual settings, leading to unsuccessful speech recognition outcomes. The present study introduces convolutional recurrent neural network (CRNN) based LID, designed to operate on the Mel-frequency Cepstral Coefficient (MFCC) characteristics of audio samples. Furthermore, we replicate certain state-of-the-art methodologies, specifically the Convolutional Neural Network (CNN) and Attention-based Convolutional Recurrent Neural Network (CRNN with attention), and conduct a comparative analysis with our CRNN-based approach. We conducted comprehensive evaluations on thirteen distinct Indian languages and our model resulted in over 98\% classification accuracy. The LID model exhibits high-performance levels ranging from 97% to 100% for languages that are linguistically similar. The proposed LID model exhibits a high degree of extensibility to additional languages and demonstrates a strong resistance to noise, achieving 91.2% accuracy in a noisy setting when applied to a European Language (EU) dataset.

PaperPDF

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)General ClassificationLanguage IdentificationSpeech RecognitionSpoken language identificationspeech-recognition

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Spoken language identification IndicTTS CRNN Classification Accuracy 0.987 #1 of 3 Archive leaderboard report
Spoken language identification IndicTTS CRNN Attention Classification Accuracy 0.987 #2 of 3 Archive leaderboard report
Spoken language identification IndicTTS CNN Classification Accuracy 0.983 #3 of 3 Archive leaderboard report
Spoken language identification YouTube News dataset (No Noise) CRNN Accuracy 0.967 #1 of 5 Archive leaderboard report
Spoken language identification YouTube News dataset (No Noise) CRNN Attention Accuracy 0.966 #2 of 5 Archive leaderboard report
Spoken language identification YouTube News dataset (No Noise) CNN Accuracy 0.948 #4 of 5 Archive leaderboard report
Spoken language identification YouTube News dataset (White Noise) CRNN Accuracy 0.912 #1 of 5 Archive leaderboard report
Spoken language identification YouTube News dataset (White Noise) CRNN Attention Accuracy 0.888 #3 of 5 Archive leaderboard report
Spoken language identification YouTube News dataset (White Noise) CNN Accuracy 0.871 #4 of 5 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections