{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/on-learning-associations-of-faces-and-voices","title":"On Learning Associations of Faces and Voices","arxiv_id":"1805.05553","date":"2018-05-15","proceeding":null,"authors":["Changil Kim","Hijung Valentina Shin","Tae-Hyun Oh","Alexandre Kaspar","Mohamed Elgharib","Wojciech Matusik"],"abstract":"In this paper, we study the associations between human faces and voices.\nAudiovisual integration, specifically the integration of facial and vocal\ninformation is a well-researched area in neuroscience. It is shown that the\noverlapping information between the two modalities plays a significant role in\nperceptual tasks such as speaker identification. Through an online study on a\nnew dataset we created, we confirm previous findings that people can associate\nunseen faces with corresponding voices and vice versa with greater than chance\naccuracy. We computationally model the overlapping information between faces\nand voices and show that the learned cross-modal representation contains enough\ninformation to identify matching faces and voices with performance similar to\nthat of humans. Our representation exhibits correlations to certain demographic\nattributes and features obtained from either visual or aural modality alone. We\nrelease our dataset of audiovisual recordings and demographic annotations of\npeople reading out short text used in our studies.","url_abs":"http://arxiv.org/abs/1805.05553v3","url_pdf":"http://arxiv.org/pdf/1805.05553v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"on-learning-associations-of-faces-and-voices","repo_url":"https://github.com/changil/facevoice","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"speaker-identification","task_name":"Speaker Identification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1805.05553","atlas_url":"https://app.syntology.ai/?focus=1805.05553","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1805.05553"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/changil/facevoice","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"unverified":3},"by_repo_kind":{"listed":{"samples":3,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"292d30c5795a85f6","entry":"bn","repo":"changil/facevoice","repo_kind":"listed","path":"facevoice.py","file_url":"https://github.com/changil/facevoice/blob/HEAD/facevoice.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"292d30c5795a85f6"}},{"code_sha256_prefix":"1bb7d4fdf72a3c03","entry":"conv","repo":"changil/facevoice","repo_kind":"listed","path":"facevoice.py","file_url":"https://github.com/changil/facevoice/blob/HEAD/facevoice.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"1bb7d4fdf72a3c03"}},{"code_sha256_prefix":"041d099ede5f732f","entry":"fc","repo":"changil/facevoice","repo_kind":"listed","path":"facevoice.py","file_url":"https://github.com/changil/facevoice/blob/HEAD/facevoice.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"041d099ede5f732f"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}