{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/diversity-and-bias-in-audio-captioning","title":"Diversity and bias in audio captioning datasets","arxiv_id":null,"date":"2022-11-15","proceeding":"DCASE workshop 2022 11","authors":["Irene Martin-Morato","Annamaria Mesaros"],"abstract":"Describing soundscapes in sentences allows better understand- ing of the acoustic scene than a single label indicating the acoustic scene class or a set of audio tags indicating the sound events active in the audio clip. In addition, the richness of natural language allows a range of possible descriptions for the same acoustic scene. In this work, we address the diversity obtained when collecting descrip- tions of soundscapes using crowdsourcing. We study how much the collection of audio captions can be guided by the instructions given in the annotation task, by analysing the possible bias introduced by auxiliary information provided in the annotation process. Our study shows that even when hints are given with the audio content, different annotators describe the same soundscape using different vocabulary. In automatic captioning, hints provided as audio tags represent grounding textual information that facilitates guiding the captioning output towards specific concepts. We also release a new dataset of audio captions and audio tags produced by multiple anno- tators for a subset of the TAU Urban Acoustic Scenes 2019 dataset, suitable for studying guided captioning.","url_abs":"https://dcase.community/workshop2021/proceedings","url_pdf":"https://dcase.community/documents/workshop2021/proceedings/DCASE2021Workshop_Martin_34.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"audio-captioning","task_name":"Audio captioning"},{"task_slug":"diversity","task_name":"Diversity"}],"methods":[],"datasets_introduced":[{"slug":"macs","name":"MACS","full_name":"Multi-Annotator Captioned Soundscapes"}],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}