{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/birdsoundsdenoising-deep-visual-audio","title":"BirdSoundsDenoising: Deep Visual Audio Denoising for Bird Sounds","arxiv_id":"2210.10196","date":"2022-10-18","proceeding":null,"authors":["Youshan Zhang","Jialu Li"],"abstract":"Audio denoising has been explored for decades using both traditional and deep learning-based methods. However, these methods are still limited to either manually added artificial noise or lower denoised audio quality. To overcome these challenges, we collect a large-scale natural noise bird sound dataset. We are the first to transfer the audio denoising problem into an image segmentation problem and propose a deep visual audio denoising (DVAD) model. With a total of 14,120 audio images, we develop an audio ImageMask tool and propose to use a few-shot generalization strategy to label these images. Extensive experimental results demonstrate that the proposed model achieves state-of-the-art performance. We also show that our method can be easily generalized to speech denoising, audio separation, audio enhancement, and noise estimation.","url_abs":"https://arxiv.org/abs/2210.10196v1","url_pdf":"https://arxiv.org/pdf/2210.10196v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"birdsoundsdenoising-deep-visual-audio","repo_url":"https://github.com/youshanzhang/birdsoundsdenoising","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"audio-denoising","task_name":"Audio Denoising"},{"task_slug":"denoising","task_name":"Denoising"},{"task_slug":"image-segmentation","task_name":"Image Segmentation"},{"task_slug":"noise-estimation","task_name":"Noise Estimation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"speech-denoising","task_name":"Speech Denoising"}],"methods":[],"datasets_introduced":[{"slug":"birdsoundsdenoising-deep-visual-audio","name":"BirdSoundsDenoising: Deep Visual Audio Denoising for Bird Sounds","full_name":""}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}