{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-speaker-representations-with-mutual","title":"Learning Speaker Representations with Mutual Information","arxiv_id":"1812.00271","date":"2018-12-01","proceeding":null,"authors":["Mirco Ravanelli","Yoshua Bengio"],"abstract":"Learning good representations is of crucial importance in deep learning.\nMutual Information (MI) or similar measures of statistical dependence are\npromising tools for learning these representations in an unsupervised way. Even\nthough the mutual information between two random variables is hard to measure\ndirectly in high dimensional spaces, some recent studies have shown that an\nimplicit optimization of MI can be achieved with an encoder-discriminator\narchitecture similar to that of Generative Adversarial Networks (GANs). In this\nwork, we learn representations that capture speaker identities by maximizing\nthe mutual information between the encoded representations of chunks of speech\nrandomly sampled from the same sentence. The proposed encoder relies on the\nSincNet architecture and transforms raw speech waveform into a compact feature\nvector. The discriminator is fed by either positive samples (of the joint\ndistribution of encoded chunks) or negative samples (from the product of the\nmarginals) and is trained to separate them. We report experiments showing that\nthis approach effectively learns useful speaker representations, leading to\npromising results on speaker identification and verification tasks. Our\nexperiments consider both unsupervised and semi-supervised settings and compare\nthe performance achieved with different objective functions.","url_abs":"http://arxiv.org/abs/1812.00271v2","url_pdf":"http://arxiv.org/pdf/1812.00271v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-speaker-representations-with-mutual","repo_url":"https://github.com/Js-Mim/rl_singing_voice","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"learning-speaker-representations-with-mutual","repo_url":"https://github.com/theolepage/ssl-for-slr","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"sentence","task_name":"Sentence"},{"task_slug":"speaker-identification","task_name":"Speaker Identification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1812.00271","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1812.00271"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/theolepage/ssl-for-slr","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/Js-Mim/rl_singing_voice","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"unverified":6},"by_repo_kind":{"listed":{"samples":6,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"b56fa43cfb65baf6","entry":"mse","repo":"Js-Mim/rl_singing_voice","repo_kind":"listed","path":"nn_modules/losses.py","file_url":"https://github.com/Js-Mim/rl_singing_voice/blob/HEAD/nn_modules/losses.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"b56fa43cfb65baf6"}},{"code_sha256_prefix":"9f635843dccf2ce0","entry":"neg_snr","repo":"Js-Mim/rl_singing_voice","repo_kind":"listed","path":"nn_modules/losses.py","file_url":"https://github.com/Js-Mim/rl_singing_voice/blob/HEAD/nn_modules/losses.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"9f635843dccf2ce0"}},{"code_sha256_prefix":"1dbf377179883c8a","entry":"si_sdr","repo":"Js-Mim/rl_singing_voice","repo_kind":"listed","path":"nn_modules/losses.py","file_url":"https://github.com/Js-Mim/rl_singing_voice/blob/HEAD/nn_modules/losses.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"1dbf377179883c8a"}},{"code_sha256_prefix":"38c98d98a1fbbfce","entry":"sinkhorn","repo":"Js-Mim/rl_singing_voice","repo_kind":"listed","path":"nn_modules/sk_iter.py","file_url":"https://github.com/Js-Mim/rl_singing_voice/blob/HEAD/nn_modules/sk_iter.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"38c98d98a1fbbfce"}},{"code_sha256_prefix":"257af754360fbf08","entry":"to_mel","repo":"Js-Mim/rl_singing_voice","repo_kind":"listed","path":"nn_modules/cls_fe_nnct.py","file_url":"https://github.com/Js-Mim/rl_singing_voice/blob/HEAD/nn_modules/cls_fe_nnct.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"257af754360fbf08"}},{"code_sha256_prefix":"f0132befe1ee9fb1","entry":"to_normalized_hz","repo":"Js-Mim/rl_singing_voice","repo_kind":"listed","path":"nn_modules/cls_fe_nnct.py","file_url":"https://github.com/Js-Mim/rl_singing_voice/blob/HEAD/nn_modules/cls_fe_nnct.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"f0132befe1ee9fb1"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}