{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-problem-agnostic-speech","title":"Learning Problem-agnostic Speech Representations from Multiple Self-supervised Tasks","arxiv_id":"1904.03416","date":"2019-04-06","proceeding":null,"authors":["Santiago Pascual","Mirco Ravanelli","Joan Serrà","Antonio Bonafonte","Yoshua Bengio"],"abstract":"Learning good representations without supervision is still an open issue in\nmachine learning, and is particularly challenging for speech signals, which are\noften characterized by long sequences with a complex hierarchical structure.\nSome recent works, however, have shown that it is possible to derive useful\nspeech representations by employing a self-supervised encoder-discriminator\napproach. This paper proposes an improved self-supervised method, where a\nsingle neural encoder is followed by multiple workers that jointly solve\ndifferent self-supervised tasks. The needed consensus across different tasks\nnaturally imposes meaningful constraints to the encoder, contributing to\ndiscover general representations and to minimize the risk of learning\nsuperficial ones. Experiments show that the proposed approach can learn\ntransferable, robust, and problem-agnostic features that carry on relevant\ninformation from the speech signal, such as speaker identity, phonemes, and\neven higher-level features such as emotional cues. In addition, a number of\ndesign choices make the encoder easily exportable, facilitating its direct\nusage or adaptation to different problems.","url_abs":"http://arxiv.org/abs/1904.03416v1","url_pdf":"http://arxiv.org/pdf/1904.03416v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-problem-agnostic-speech","repo_url":"https://github.com/santi-pdp/pase","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"distant-speech-recognition","task_name":"Distant Speech Recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/distant-speech-recognition-on-dirha-english","task":"Distant Speech Recognition","dataset":"DIRHA English WSJ","model":"PASE-FineTuned","rank_in_archive_order":2,"of":3,"metrics":{"Word Error Rate (WER)":"29.8"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1904.03416","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}