{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/visual-to-sound-generating-natural-sound-for","title":"Visual to Sound: Generating Natural Sound for Videos in the Wild","arxiv_id":"1712.01393","date":"2017-12-04","proceeding":"CVPR 2018 6","authors":["Yipin Zhou","Zhaowen Wang","Chen Fang","Trung Bui","Tamara L. Berg"],"abstract":"As two of the five traditional human senses (sight, hearing, taste, smell,\nand touch), vision and sound are basic sources through which humans understand\nthe world. Often correlated during natural events, these two modalities combine\nto jointly affect human perception. In this paper, we pose the task of\ngenerating sound given visual input. Such capabilities could help enable\napplications in virtual reality (generating sound for virtual scenes\nautomatically) or provide additional accessibility to images or videos for\npeople with visual impairments. As a first step in this direction, we apply\nlearning-based methods to generate raw waveform samples given input video\nframes. We evaluate our models on a dataset of videos containing a variety of\nsounds (such as ambient sounds and sounds from people/animals). Our experiments\nshow that the generated sounds are fairly realistic and have good temporal\nsynchronization with the visual inputs.","url_abs":"http://arxiv.org/abs/1712.01393v2","url_pdf":"http://arxiv.org/pdf/1712.01393v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"visual-to-sound-generating-natural-sound-for","repo_url":"https://github.com/PeihaoChen/regnet","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"visual-to-sound-generating-natural-sound-for","repo_url":"https://github.com/ZenzenDatabase/CMR_audiovisual","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}},{"paper_slug":"visual-to-sound-generating-natural-sound-for","repo_url":"https://github.com/dddzeng/CMR_audiovisual","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1712.01393","atlas_url":"https://app.syntology.ai/?focus=1712.01393","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}