{"url":"/dataset/is3-interactive-synthetic-sound-source","name":"IS3 (Interactive-Synthetic Sound Source) Dataset","full_name":null,"description_markdown":"We introduce a new synthetic test set named IS3 for interactive sound source localization. By leveraging diffusion models, we generate images containing multiple sounding objects. Any\r\ncombination of sounding objects can appear in the same scene. Additionally, this dataset offers unusual scenes and unique combinations that are rarely found in nature, such as ‘a donkey playing a saxophone’ or ‘a sea lion on the snow’. This dataset provides both segmentation maps and bounding box information with class categories. IS3 includes 3240 images, resulting in 6480 unique audio-visual instances (with 2 objects per image) across 118 categories.\r\nThis dataset can be used in below tasks:\r\n1) Sound Source Localization\r\n2) Audio-Visual Segmentation\r\n3) Semantic Segmentation","description_withheld":null,"homepage":"https://github.com/kaistmm/SSLalignment","introduced_date":"2024-07-18","introduced_date_note":null,"introduced_by":{"paper":"/paper/aligning-sight-and-sound-advanced-sound","title":"Aligning Sight and Sound: Advanced Sound Source Localization Through Audio-Visual Alignment","first_author":"Arda Senocak","url":null},"license":null,"modalities":[{"name":"Images","url":"/datasets/modality/images"},{"name":"Audio","url":"/datasets/modality/audio"}],"tasks":[{"name":"audio-visual learning","url":"/task/audio-visual-learning","datasets_with_task":"/datasets/task/audio-visual-learning"}],"languages":[],"variants":["IS3 (Interactive-Synthetic Sound Source) Dataset"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}