Papers › Spoken Language Identification using ConvNets

Spoken Language Identification using ConvNets

9 Oct 2019arXiv:1910.04269archive 2025-07-28

Sarthak, Shikhar Shukla, Govind Mittal

Language Identification (LI) is an important first step in several speech processing systems. With a growing number of voice-based assistants, speech LI has emerged as a widely researched field. To approach the problem of identifying languages, we can either adopt an implicit approach where only the speech for a language is present or an explicit one where text is available with its corresponding transcript. This paper focuses on an implicit approach due to the absence of transcriptive data. This paper benchmarks existing models and proposes a new attention based model for language identification which uses log-Mel spectrogram images as input. We also present the effectiveness of raw waveforms as features to neural network models for LI tasks. For training and evaluation of models, we classified six languages (English, French, German, Spanish, Russian and Italian) with an accuracy of 95.4% and four languages (English, French, German, Spanish) with an accuracy of 96.3% obtained from the VoxForge dataset. This approach can further be scaled to incorporate more languages.

PaperPDF

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Keyword SpottingLanguage IdentificationSpoken language identification

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Keyword Spotting VoxForge 1D-ConvNet Accuracy (%) 93.7 #1 of 2 Archive leaderboard report
Keyword Spotting VoxForge 2D-ConvNet Accuracy (%) 95.4 #2 of 2 Archive leaderboard report
Spoken language identification VoxForge Commonwealth 2D ConvNet(MixUp=YES) Accuracy (%) 95.4 #1 of 4 Archive leaderboard report
Spoken language identification VoxForge Commonwealth 2D ConvNet with Attention and GRU(MixUp=YES) Accuracy (%) 95.0 #2 of 4 Archive leaderboard report
Spoken language identification VoxForge Commonwealth 2D ConvNet(MixUp=NO) Accuracy (%) 94.3 #3 of 4 Archive leaderboard report
Spoken language identification VoxForge Commonwealth 1D ConvNet(MixUp=NO) Accuracy (%) 93.7 #4 of 4 Archive leaderboard report
Spoken language identification VoxForge European 2D ConvNet(MixUp=YES) Accuracy (%) 96.3 #1 of 5 Archive leaderboard report
Spoken language identification VoxForge European 2D ConvNet(MixUp=NO) Accuracy (%) 96.0 #2 of 5 Archive leaderboard report
Spoken language identification VoxForge European 2D ConvNet with Attention and GRU(MixUp=NO) Accuracy (%) 94.7 #3 of 5 Archive leaderboard report
Spoken language identification VoxForge European 1D ConvNet(MixUp=NO) Accuracy (%) 94.4 #4 of 5 Archive leaderboard report
Spoken language identification VoxForge European 2D ConvNet with Attention and GRU(MixUp=YES) Accuracy (%) 93.7 #5 of 5 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections