Datasets › IS3 (Interactive-Synthetic Sound Source) Dataset

IS3 (Interactive-Synthetic Sound Source) Dataset

Introduced by Arda Senocak et al. in Aligning Sight and Sound: Advanced Sound Source Localization Through Audio-Visual Alignment18 Jul 2024 archive 2025-07-28

We introduce a new synthetic test set named IS3 for interactive sound source localization. By leveraging diffusion models, we generate images containing multiple sounding objects. Any combination of sounding objects can appear in the same scene. Additionally, this dataset offers unusual scenes and unique combinations that are rarely found in nature, such as ‘a donkey playing a saxophone’ or ‘a sea lion on the snow’. This dataset provides both segmentation maps and bounding box information with class categories. IS3 includes 3240 images, resulting in 6480 unique audio-visual instances (with 2 objects per image) across 118 categories. This dataset can be used in below tasks: 1) Sound Source Localization 2) Audio-Visual Segmentation 3) Semantic Segmentation

Benchmarks archive 2025-07-28

No leaderboard in the archive resolves to this dataset.

Papers archive 2025-07-28

No paper in the archive has a leaderboard row on this dataset; the archive counts 1 paper for it but never published that list.

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • IS3 (Interactive-Synthetic Sound Source) Dataset

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections