{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/audio-retrieval-with-natural-language-queries-1","title":"Audio Retrieval with Natural Language Queries: A Benchmark Study","arxiv_id":"2112.09418","date":"2021-12-17","proceeding":null,"authors":["A. Sophia Koepke","Andreea-Maria Oncescu","João F. Henriques","Zeynep Akata","Samuel Albanie"],"abstract":"The objectives of this work are cross-modal text-audio and audio-text retrieval, in which the goal is to retrieve the audio content from a pool of candidates that best matches a given written description and vice versa. Text-audio retrieval enables users to search large databases through an intuitive interface: they simply issue free-form natural language descriptions of the sound they would like to hear. To study the tasks of text-audio and audio-text retrieval, which have received limited attention in the existing literature, we introduce three challenging new benchmarks. We first construct text-audio and audio-text retrieval benchmarks from the AudioCaps and Clotho audio captioning datasets. Additionally, we introduce the SoundDescs benchmark, which consists of paired audio and natural language descriptions for a diverse collection of sounds that are complementary to those found in AudioCaps and Clotho. We employ these three benchmarks to establish baselines for cross-modal text-audio and audio-text retrieval, where we demonstrate the benefits of pre-training on diverse audio tasks. We hope that our benchmarks will inspire further research into audio retrieval with free-form text queries. Code, audio features for all datasets used, and the SoundDescs dataset are publicly available at https://github.com/akoepke/audio-retrieval-benchmark.","url_abs":"https://arxiv.org/abs/2112.09418v2","url_pdf":"https://arxiv.org/pdf/2112.09418v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"audio-retrieval-with-natural-language-queries-1","repo_url":"https://github.com/akoepke/audio-retrieval-benchmark","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"audio-captioning","task_name":"Audio captioning"},{"task_slug":"audio-to-text-retrieval","task_name":"Audio to Text Retrieval"},{"task_slug":null,"task_name":"AudioCaps"},{"task_slug":"natural-language-queries","task_name":"Natural Language Queries"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"text-retrieval","task_name":"Text Retrieval"},{"task_slug":"text-to-audio-retrieval","task_name":"Text to Audio Retrieval"}],"methods":[],"datasets_introduced":[{"slug":"sounddescs","name":"SoundDescs","full_name":""}],"methods_introduced":[],"results":[{"leaderboard":"/sota/text-to-audio-retrieval-on-audiocaps","task":"Text to Audio Retrieval","dataset":"AudioCaps","model":"MMT","rank_in_archive_order":5,"of":11,"metrics":{"R@1":"36.1±3.3","R@10":"84.5±2.0"},"uses_additional_data":false},{"leaderboard":"/sota/text-to-audio-retrieval-on-audiocaps","task":"Text to Audio Retrieval","dataset":"AudioCaps","model":"CE","rank_in_archive_order":8,"of":11,"metrics":{"R@1":"23.6± 0.6","R@10":"71.4±0.5"},"uses_additional_data":false},{"leaderboard":"/sota/text-to-audio-retrieval-on-audiocaps","task":"Text to Audio Retrieval","dataset":"AudioCaps","model":"MoEE","rank_in_archive_order":10,"of":11,"metrics":{"R@1":"23.0±0.7","R@10":"71.0±1.2"},"uses_additional_data":false},{"leaderboard":"/sota/text-to-audio-retrieval-on-clotho","task":"Text to Audio Retrieval","dataset":"Clotho","model":"MMT","rank_in_archive_order":10,"of":12,"metrics":{"R@1":"6.5±0.6","R@10":"32.8±2.1"},"uses_additional_data":false},{"leaderboard":"/sota/text-to-audio-retrieval-on-clotho","task":"Text to Audio Retrieval","dataset":"Clotho","model":"CE(pretraining:SoundDescs)","rank_in_archive_order":11,"of":12,"metrics":{"R@1":"6.4±0.5","R@10":"32.5±1.7"},"uses_additional_data":false},{"leaderboard":"/sota/text-to-audio-retrieval-on-sounddescs","task":"Text to Audio Retrieval","dataset":"SoundDescs","model":"CE","rank_in_archive_order":1,"of":4,"metrics":{"R@1":"31.1±0.2","R@10":"70.8±0.5"},"uses_additional_data":false},{"leaderboard":"/sota/text-to-audio-retrieval-on-sounddescs","task":"Text to Audio Retrieval","dataset":"SoundDescs","model":"MoEE","rank_in_archive_order":2,"of":4,"metrics":{"R@1":"30.8±0.7","R@10":"70.9±0.5"},"uses_additional_data":false},{"leaderboard":"/sota/text-to-audio-retrieval-on-sounddescs","task":"Text to Audio Retrieval","dataset":"SoundDescs","model":"MMT","rank_in_archive_order":3,"of":4,"metrics":{"R@1":"30.7±0.4","R@10":"72.7±0.8"},"uses_additional_data":false},{"leaderboard":"/sota/text-to-audio-retrieval-on-sounddescs","task":"Text to Audio Retrieval","dataset":"SoundDescs","model":"CE(pretrained: AudioCaps)","rank_in_archive_order":4,"of":4,"metrics":{"R@1":"23.3±0.7","R@10":"63.9±0.5"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/2112.09418","atlas_url":"https://app.syntology.ai/?focus=2112.09418","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}