{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/25d-visual-sound","title":"2.5D Visual Sound","arxiv_id":"1812.04204","date":"2018-12-11","proceeding":"CVPR 2019 6","authors":["Ruohan Gao","Kristen Grauman"],"abstract":"Binaural audio provides a listener with 3D sound sensation, allowing a rich\nperceptual experience of the scene. However, binaural recordings are scarcely\navailable and require nontrivial expertise and equipment to obtain. We propose\nto convert common monaural audio into binaural audio by leveraging video. The\nkey idea is that visual frames reveal significant spatial cues that, while\nexplicitly lacking in the accompanying single-channel audio, are strongly\nlinked to it. Our multi-modal approach recovers this link from unlabeled video.\nWe devise a deep convolutional neural network that learns to decode the\nmonaural (single-channel) soundtrack into its binaural counterpart by injecting\nvisual information about object and scene configurations. We call the resulting\noutput 2.5D visual sound---the visual stream helps \"lift\" the flat single\nchannel audio into spatialized sound. In addition to sound generation, we show\nthe self-supervised representation learned by our network benefits audio-visual\nsource separation. Our video results:\nhttp://vision.cs.utexas.edu/projects/2.5D_visual_sound/","url_abs":"http://arxiv.org/abs/1812.04204v4","url_pdf":"http://arxiv.org/pdf/1812.04204v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"25d-visual-sound","repo_url":"https://github.com/facebookresearch/FAIR-Play","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"CC-BY-4.0"}},{"paper_slug":"25d-visual-sound","repo_url":"https://github.com/facebookresearch/2.5D-Visual-Sound","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"CC-BY-4.0"}}],"tasks":[],"methods":[],"datasets_introduced":[{"slug":"fair-play","name":"FAIR-Play","full_name":"FAIR-Play"}],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1812.04204","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}