{"url":"/dataset/vgg-ss","name":"VGG-SS","full_name":"VGG-Sound Source","description_markdown":"VGG-SS (VGG Sound Source) is a benchmark for evaluating sound source localisation in videos. The dataset consists on a new set of annotations for the recently-introduced [VGG-Sound dataset](vgg-sound), where the sound sources visible in each video clip are explicitly marked with bounding box annotations. This dataset is 20 times larger than analogous existing ones, contains 5K videos spanning over 200 categories, and, differently from Flickr SoundNet, is video-based.","description_withheld":null,"homepage":"https://www.robots.ox.ac.uk/~vgg/research/lvs/","introduced_date":"2021-04-06","introduced_date_note":null,"introduced_by":{"paper":"/paper/localizing-visual-sounds-the-hard-way","title":"Localizing Visual Sounds the Hard Way","first_author":"Honglie Chen","url":null},"license":{"name":"Unknown","url":null},"modalities":[{"name":"Videos","url":"/datasets/modality/videos"}],"tasks":[],"languages":[],"variants":["VGG-SS"],"data_loaders":[],"num_papers_in_archive":30,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}