{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/spoken-language-identification-using-convnets","title":"Spoken Language Identification using ConvNets","arxiv_id":"1910.04269","date":"2019-10-09","proceeding":null,"authors":["Sarthak","Shikhar Shukla","Govind Mittal"],"abstract":"Language Identification (LI) is an important first step in several speech processing systems. With a growing number of voice-based assistants, speech LI has emerged as a widely researched field. To approach the problem of identifying languages, we can either adopt an implicit approach where only the speech for a language is present or an explicit one where text is available with its corresponding transcript. This paper focuses on an implicit approach due to the absence of transcriptive data. This paper benchmarks existing models and proposes a new attention based model for language identification which uses log-Mel spectrogram images as input. We also present the effectiveness of raw waveforms as features to neural network models for LI tasks. For training and evaluation of models, we classified six languages (English, French, German, Spanish, Russian and Italian) with an accuracy of 95.4% and four languages (English, French, German, Spanish) with an accuracy of 96.3% obtained from the VoxForge dataset. This approach can further be scaled to incorporate more languages.","url_abs":"https://arxiv.org/abs/1910.04269v1","url_pdf":"https://arxiv.org/pdf/1910.04269v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"keyword-spotting","task_name":"Keyword Spotting"},{"task_slug":"language-identification","task_name":"Language Identification"},{"task_slug":"spoken-language-identification","task_name":"Spoken language identification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/keyword-spotting-on-voxforge","task":"Keyword Spotting","dataset":"VoxForge","model":"1D-ConvNet","rank_in_archive_order":1,"of":2,"metrics":{"Accuracy (%)":"93.7"},"uses_additional_data":false},{"leaderboard":"/sota/keyword-spotting-on-voxforge","task":"Keyword Spotting","dataset":"VoxForge","model":"2D-ConvNet","rank_in_archive_order":2,"of":2,"metrics":{"Accuracy (%)":"95.4"},"uses_additional_data":false},{"leaderboard":"/sota/spoken-language-identification-on-voxforge","task":"Spoken language identification","dataset":"VoxForge Commonwealth","model":"2D ConvNet(MixUp=YES)","rank_in_archive_order":1,"of":4,"metrics":{"Accuracy (%)":"95.4"},"uses_additional_data":false},{"leaderboard":"/sota/spoken-language-identification-on-voxforge","task":"Spoken language identification","dataset":"VoxForge Commonwealth","model":"2D ConvNet with Attention and GRU(MixUp=YES)","rank_in_archive_order":2,"of":4,"metrics":{"Accuracy (%)":"95.0"},"uses_additional_data":false},{"leaderboard":"/sota/spoken-language-identification-on-voxforge","task":"Spoken language identification","dataset":"VoxForge Commonwealth","model":"2D ConvNet(MixUp=NO)","rank_in_archive_order":3,"of":4,"metrics":{"Accuracy (%)":"94.3"},"uses_additional_data":false},{"leaderboard":"/sota/spoken-language-identification-on-voxforge","task":"Spoken language identification","dataset":"VoxForge Commonwealth","model":"1D ConvNet(MixUp=NO)","rank_in_archive_order":4,"of":4,"metrics":{"Accuracy (%)":"93.7"},"uses_additional_data":false},{"leaderboard":"/sota/spoken-language-identification-on-voxforge-1","task":"Spoken language identification","dataset":"VoxForge European","model":"2D ConvNet(MixUp=YES)","rank_in_archive_order":1,"of":5,"metrics":{"Accuracy (%)":"96.3"},"uses_additional_data":false},{"leaderboard":"/sota/spoken-language-identification-on-voxforge-1","task":"Spoken language identification","dataset":"VoxForge European","model":"2D ConvNet(MixUp=NO)","rank_in_archive_order":2,"of":5,"metrics":{"Accuracy (%)":"96.0"},"uses_additional_data":false},{"leaderboard":"/sota/spoken-language-identification-on-voxforge-1","task":"Spoken language identification","dataset":"VoxForge European","model":"2D ConvNet with Attention and GRU(MixUp=NO)","rank_in_archive_order":3,"of":5,"metrics":{"Accuracy (%)":"94.7"},"uses_additional_data":false},{"leaderboard":"/sota/spoken-language-identification-on-voxforge-1","task":"Spoken language identification","dataset":"VoxForge European","model":"1D ConvNet(MixUp=NO)","rank_in_archive_order":4,"of":5,"metrics":{"Accuracy (%)":"94.4"},"uses_additional_data":false},{"leaderboard":"/sota/spoken-language-identification-on-voxforge-1","task":"Spoken language identification","dataset":"VoxForge European","model":"2D ConvNet with Attention and GRU(MixUp=YES)","rank_in_archive_order":5,"of":5,"metrics":{"Accuracy (%)":"93.7"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}