{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/look-listen-and-learn","title":"Look, Listen and Learn","arxiv_id":"1705.08168","date":"2017-05-23","proceeding":"ICCV 2017 10","authors":["Relja Arandjelović","Andrew Zisserman"],"abstract":"We consider the question: what can be learnt by looking at and listening to a\nlarge number of unlabelled videos? There is a valuable, but so far untapped,\nsource of information contained in the video itself -- the correspondence\nbetween the visual and the audio streams, and we introduce a novel\n\"Audio-Visual Correspondence\" learning task that makes use of this. Training\nvisual and audio networks from scratch, without any additional supervision\nother than the raw unconstrained videos themselves, is shown to successfully\nsolve this task, and, more interestingly, result in good visual and audio\nrepresentations. These features set the new state-of-the-art on two sound\nclassification benchmarks, and perform on par with the state-of-the-art\nself-supervised approaches on ImageNet classification. We also demonstrate that\nthe network is able to localize objects in both modalities, as well as perform\nfine-grained recognition tasks.","url_abs":"http://arxiv.org/abs/1705.08168v2","url_pdf":"http://arxiv.org/pdf/1705.08168v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"look-listen-and-learn","repo_url":"https://github.com/marl/l3embedding","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"audio-classification","task_name":"Audio Classification"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"sound-classification","task_name":"Sound Classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/audio-classification-on-audioset","task":"Audio Classification","dataset":"AudioSet","model":"L3","rank_in_archive_order":50,"of":51,"metrics":{"Test mAP":"0.249"},"uses_additional_data":false},{"leaderboard":"/sota/audio-classification-on-esc-50","task":"Audio Classification","dataset":"ESC-50","model":"L3","rank_in_archive_order":29,"of":29,"metrics":{"Top-1 Accuracy":"79.3"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1705.08168","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}