{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/points2sound-from-mono-to-binaural-audio","title":"Points2Sound: From mono to binaural audio using 3D point cloud scenes","arxiv_id":"2104.12462","date":"2021-04-26","proceeding":null,"authors":["Francesc Lluís","Vasileios Chatziioannou","Alex Hofmann"],"abstract":"For immersive applications, the generation of binaural sound that matches its visual counterpart is crucial to bring meaningful experiences to people in a virtual environment. Recent studies have shown the possibility of using neural networks for synthesizing binaural audio from mono audio by using 2D visual information as guidance. Extending this approach by guiding the audio with 3D visual information and operating in the waveform domain may allow for a more accurate auralization of a virtual audio scene. We propose Points2Sound, a multi-modal deep learning model which generates a binaural version from mono audio using 3D point cloud scenes. Specifically, Points2Sound consists of a vision network and an audio network. The vision network uses 3D sparse convolutions to extract a visual feature from the point cloud scene. Then, the visual feature conditions the audio network, which operates in the waveform domain, to synthesize the binaural version. Results show that 3D visual information can successfully guide multi-modal deep learning models for the task of binaural synthesis. We also investigate how 3D point cloud attributes, learning objectives, different reverberant conditions, and several types of mono mixture signals affect the binaural audio synthesis performance of Points2Sound for the different numbers of sound sources present in the scene.","url_abs":"https://arxiv.org/abs/2104.12462v3","url_pdf":"https://arxiv.org/pdf/2104.12462v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"points2sound-from-mono-to-binaural-audio","repo_url":"https://github.com/francesclluis/points2sound","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"audio-synthesis","task_name":"Audio Synthesis"}],"methods":[{"method_slug":"sparse-convolutions","method_name":"Sparse Convolutions"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2104.12462","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2104.12462"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/francesclluis/points2sound","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran":2,"unverified":2},"by_repo_kind":{"official":{"samples":4,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"c26dd101957c5772","entry":"unet_conv","repo":"francesclluis/points2sound","repo_kind":"official","path":"models/audio_net.py","file_url":"https://github.com/francesclluis/points2sound/blob/HEAD/models/audio_net.py","link_basis":"plan_row","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"c26dd101957c5772"}},{"code_sha256_prefix":"e1be3fb442f83f7f","entry":"unet_upconv","repo":"francesclluis/points2sound","repo_kind":"official","path":"models/audio_net.py","file_url":"https://github.com/francesclluis/points2sound/blob/HEAD/models/audio_net.py","link_basis":"plan_row","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"e1be3fb442f83f7f"}},{"code_sha256_prefix":"6dd7d61d18bfaea5","entry":"Envelope_distance","repo":"francesclluis/points2sound","repo_kind":"official","path":"metrics_binaural.py","file_url":"https://github.com/francesclluis/points2sound/blob/HEAD/metrics_binaural.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"6dd7d61d18bfaea5"}},{"code_sha256_prefix":"03ce5ded167a28fd","entry":"center_trim","repo":"francesclluis/points2sound","repo_kind":"official","path":"models/audio_net.py","file_url":"https://github.com/francesclluis/points2sound/blob/HEAD/models/audio_net.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"03ce5ded167a28fd"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}