{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-to-localize-sound-sources-in-visual","title":"Learning to Localize Sound Sources in Visual Scenes: Analysis and Applications","arxiv_id":"1911.09649","date":"2019-11-20","proceeding":null,"authors":["Arda Senocak","Tae-Hyun Oh","Junsik Kim","Ming-Hsuan Yang","In So Kweon"],"abstract":"Visual events are usually accompanied by sounds in our daily lives. However, can the machines learn to correlate the visual scene and sound, as well as localize the sound source only by observing them like humans? To investigate its empirical learnability, in this work we first present a novel unsupervised algorithm to address the problem of localizing sound sources in visual scenes. In order to achieve this goal, a two-stream network structure which handles each modality with attention mechanism is developed for sound source localization. The network naturally reveals the localized response in the scene without human annotation. In addition, a new sound source dataset is developed for performance evaluation. Nevertheless, our empirical evaluation shows that the unsupervised method generates false conclusions in some cases. Thereby, we show that this false conclusion cannot be fixed without human prior knowledge due to the well-known correlation and causality mismatch misconception. To fix this issue, we extend our network to the supervised and semi-supervised network settings via a simple modification due to the general architecture of our two-stream network. We show that the false conclusions can be effectively corrected even with a small amount of supervision, i.e., semi-supervised setup. Furthermore, we present the versatility of the learned audio and visual embeddings on the cross-modal content alignment and we extend this proposed algorithm to a new application, sound saliency based automatic camera view panning in 360-degree{\\deg} videos.","url_abs":"https://arxiv.org/abs/1911.09649v1","url_pdf":"https://arxiv.org/pdf/1911.09649v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-to-localize-sound-sources-in-visual","repo_url":"https://github.com/ardasnck/learning_to_localize_sound_source","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"sound-source-localization","task_name":"Sound Source Localization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1911.09649","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1911.09649"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/ardasnck/learning_to_localize_sound_source","reach":null}],"summary":{"ran_honours":2,"ran_draft_wrong":2},"by_repo_kind":{"listed":{"samples":4,"ran":4,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":4,"samples":[{"code_sha256_prefix":"debf0bb800e2bdcf","entry":"audio_loader","repo":"ardasnck/learning_to_localize_sound_source","repo_kind":"listed","path":"Sound_Localization_Dataset.py","file_url":"https://github.com/ardasnck/learning_to_localize_sound_source/blob/HEAD/Sound_Localization_Dataset.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"debf0bb800e2bdcf"}},{"code_sha256_prefix":"72d6dd2fcebb2346","entry":"image_loader","repo":"ardasnck/learning_to_localize_sound_source","repo_kind":"listed","path":"Sound_Localization_Dataset.py","file_url":"https://github.com/ardasnck/learning_to_localize_sound_source/blob/HEAD/Sound_Localization_Dataset.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"72d6dd2fcebb2346"}},{"code_sha256_prefix":"bbcd81d386c6b590","entry":"localization_gt_loader","repo":"ardasnck/learning_to_localize_sound_source","repo_kind":"listed","path":"Sound_Localization_Dataset.py","file_url":"https://github.com/ardasnck/learning_to_localize_sound_source/blob/HEAD/Sound_Localization_Dataset.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"bbcd81d386c6b590"}},{"code_sha256_prefix":"660490c420fbdf3e","entry":"overlay","repo":"ardasnck/learning_to_localize_sound_source","repo_kind":"listed","path":"sound_localization_main.py","file_url":"https://github.com/ardasnck/learning_to_localize_sound_source/blob/HEAD/sound_localization_main.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":"invariant","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"660490c420fbdf3e"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}