{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/speech-vgg-a-deep-feature-extractor-for","title":"Word-level Embeddings for Cross-Task Transfer Learning in Speech Processing","arxiv_id":"1910.09909","date":"2019-10-22","proceeding":null,"authors":["Pierre Beckmann","Mikolaj Kegler","Milos Cernak"],"abstract":"Recent breakthroughs in deep learning often rely on representation learning and knowledge transfer. In recent years, unsupervised and self-supervised techniques for learning speech representation were developed to foster automatic speech recognition. Up to date, most of these approaches are task-specific and designed for within-task transfer learning between different datasets or setups of a particular task. In turn, learning task-independent representation of speech and cross-task applications of transfer learning remain less common. Here, we introduce an encoder capturing word-level representations of speech for cross-task transfer learning. We demonstrate the application of the pre-trained encoder in four distinct speech and audio processing tasks: (i) speech enhancement, (ii) language identification, (iii) speech, noise, and music classification, and (iv) speaker identification. In each task, we compare the performance of our cross-task transfer learning approach to task-specific baselines. Our results show that the speech representation captured by the encoder through the pre-training is transferable across distinct speech processing tasks and datasets. Notably, even simple applications of our pre-trained encoder outperformed task-specific methods, or were comparable, depending on the task.","url_abs":"https://arxiv.org/abs/1910.09909v5","url_pdf":"https://arxiv.org/pdf/1910.09909v5.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"speech-vgg-a-deep-feature-extractor-for","repo_url":"https://github.com/bepierre/SpeechVGG","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"GPL-3.0"}},{"paper_slug":"speech-vgg-a-deep-feature-extractor-for","repo_url":"https://github.com/MKegler/SpeechVGG","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"automatic-speech-recognition-2","task_name":"Automatic Speech Recognition"},{"task_slug":"automatic-speech-recognition","task_name":"Automatic Speech Recognition (ASR)"},{"task_slug":"language-identification","task_name":"Language Identification"},{"task_slug":"music-classification","task_name":"Music Classification"},{"task_slug":"representation-learning","task_name":"Representation Learning"},{"task_slug":"speaker-identification","task_name":"Speaker Identification"},{"task_slug":"speech-enhancement","task_name":"Speech Enhancement"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"transfer-learning","task_name":"Transfer Learning"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1910.09909","atlas_url":"https://app.syntology.ai/?focus=1910.09909","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}