{"url":"/dataset/music-avqa-r","name":"MUSIC-AVQA-R","full_name":null,"description_markdown":"We introduce the first dataset, ***MUSIC-AVQA-R***, to evaluate the robustness of AVQA models. The construction of this dataset involves two key processes: *rephrasing* and *splitting*. The former involves the rephrasing of questions in the test split of [***MUSIC-AVQA***](https://github.com/GeWu-Lab/MUSIC-AVQA), and the latter is dedicated to the categorization of questions into frequent (head) and rare (tail) subset. \r\n\r\nWe followed the previous work in partitioning the dataset into \"head\" and \"tail\" categories. Based on the number of answers in the dataset, answers with a count greater than $1.2$ times the mean, denoted as $\\mu(a)$, were categorized as \"head\" while those with counts less than  $1.2\\mu(a)$ were categorized as \"tail\"","description_withheld":null,"homepage":"https://github.com/reml-group/music-avqa-r","introduced_date":"2024-04-18","introduced_date_note":null,"introduced_by":{"paper":"/paper/look-listen-and-answer-overcoming-biases-for","title":"Look, Listen, and Answer: Overcoming Biases for Audio-Visual Question Answering","first_author":"Jie Ma","url":null},"license":null,"modalities":[],"tasks":[{"name":"Audio-Video Question Answering (AVQA)","url":"/task/audio-video-question-answering-avqa","datasets_with_task":"/datasets/task/audio-video-question-answering-avqa"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["MUSIC-AVQA-R"],"data_loaders":[],"num_papers_in_archive":2,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}