Papers › Co-Separating Sounds of Visual Objects
Co-Separating Sounds of Visual Objects
Ruohan Gao, Kristen Grauman
Learning how objects sound from video is challenging, since they often heavily overlap in a single audio channel. Current methods for visually-guided audio source separation sidestep the issue by training with artificially mixed video clips, but this puts unwieldy restrictions on training data collection and may even prevent learning the properties of "true" mixed sounds. We introduce a co-separation training paradigm that permits learning object-level sounds from unlabeled multi-source videos. Our novel training objective requires that the deep neural network's separated audio for similar-looking objects be consistently identifiable, while simultaneously reproducing accurate video-level audio tracks for each source training pair. Our approach disentangles sounds in realistic test videos, even in cases where an object was not observed individually during training. We obtain state-of-the-art results on visually-guided audio source separation and audio denoising for the MUSIC, AudioSet, and AV-Bench datasets.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Audio Denoising | AV-Bench - Guitar Solo | Co-Separation | NSDR | 11.9 | #1 of 1 | Archive leaderboard | report |
| Audio Denoising | AV-Bench - Violin Yanni | Co-Separation | NSDR | 8.53 | #1 of 1 | Archive leaderboard | report |
| Audio Denoising | AV-Bench - Wooden Horse | Co-Separation | NSDR | 14.5 | #1 of 1 | Archive leaderboard | report |
| Audio Source Separation | AudioSet | Co-Separation | SAR | 13 | #2 of 2 | Archive leaderboard | report |
| Audio Source Separation | AudioSet | Co-Separation | SDR | 4.26 | #2 of 2 | Archive leaderboard | report |
| Audio Source Separation | AudioSet | Co-Separation | SIR | 7.07 | #2 of 2 | Archive leaderboard | report |
| Audio Source Separation | MUSIC (multi-source) | Co-Separation | SAR | 11.3 | #1 of 1 | Archive leaderboard | report |
| Audio Source Separation | MUSIC (multi-source) | Co-Separation | SIR | 13.8 | #1 of 1 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections