{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/jointly-discovering-visual-objects-and-spoken","title":"Jointly Discovering Visual Objects and Spoken Words from Raw Sensory Input","arxiv_id":"1804.01452","date":"2018-04-04","proceeding":"ECCV 2018 9","authors":["David Harwath","Adrià Recasens","Dídac Surís","Galen Chuang","Antonio Torralba","James Glass"],"abstract":"In this paper, we explore neural network models that learn to associate\nsegments of spoken audio captions with the semantically relevant portions of\nnatural images that they refer to. We demonstrate that these audio-visual\nassociative localizations emerge from network-internal representations learned\nas a by-product of training to perform an image-audio retrieval task. Our\nmodels operate directly on the image pixels and speech waveform, and do not\nrely on any conventional supervision in the form of labels, segmentations, or\nalignments between the modalities during training. We perform analysis using\nthe Places 205 and ADE20k datasets demonstrating that our models implicitly\nlearn semantically-coupled object and word detectors.","url_abs":"http://arxiv.org/abs/1804.01452v1","url_pdf":"http://arxiv.org/pdf/1804.01452v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"sound-prompted-semantic-segmentation","task_name":"Sound Prompted Semantic Segmentation"},{"task_slug":"speech-prompted-semantic-segmentation","task_name":"Speech Prompted Semantic Segmentation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/sound-prompted-semantic-segmentation-on","task":"Sound Prompted Semantic Segmentation","dataset":"ADE20K","model":"DAVENet","rank_in_archive_order":4,"of":4,"metrics":{"mAP":"16.8","mIoU":"18.1"},"uses_additional_data":false},{"leaderboard":"/sota/speech-prompted-semantic-segmentation-on","task":"Speech Prompted Semantic Segmentation","dataset":"ADE20K","model":"DAVENet","rank_in_archive_order":2,"of":4,"metrics":{"mAP":"32.2","mIoU":"26.3"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1804.01452","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}