Papers › Audio Retrieval with Natural Language Queries: A Benchmark Study

Audio Retrieval with Natural Language Queries: A Benchmark Study

17 Dec 2021arXiv:2112.09418archive 2025-07-28

A. Sophia Koepke, Andreea-Maria Oncescu, João F. Henriques, Zeynep Akata, Samuel Albanie

The objectives of this work are cross-modal text-audio and audio-text retrieval, in which the goal is to retrieve the audio content from a pool of candidates that best matches a given written description and vice versa. Text-audio retrieval enables users to search large databases through an intuitive interface: they simply issue free-form natural language descriptions of the sound they would like to hear. To study the tasks of text-audio and audio-text retrieval, which have received limited attention in the existing literature, we introduce three challenging new benchmarks. We first construct text-audio and audio-text retrieval benchmarks from the AudioCaps and Clotho audio captioning datasets. Additionally, we introduce the SoundDescs benchmark, which consists of paired audio and natural language descriptions for a diverse collection of sounds that are complementary to those found in AudioCaps and Clotho. We employ these three benchmarks to establish baselines for cross-modal text-audio and audio-text retrieval, where we demonstrate the benefits of pre-training on diverse audio tasks. We hope that our benchmarks will inspire further research into audio retrieval with free-form text queries. Code, audio features for all datasets used, and the SoundDescs dataset are publicly available at https://github.com/akoepke/audio-retrieval-benchmark.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

akoepke/audio-retrieval-benchmark officialmentioned in papermentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Audio captioningAudio to Text RetrievalNatural Language QueriesRetrievalText RetrievalText to Audio Retrieval

1 archive task tag without a task page not shown.

Datasets

Introduced by this paper, per the archive.

SoundDescs

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Text to Audio Retrieval AudioCaps MMT R@1 36.1±3.3 #5 of 11 Archive leaderboard report
Text to Audio Retrieval AudioCaps MMT R@10 84.5±2.0 #5 of 11 Archive leaderboard report
Text to Audio Retrieval AudioCaps CE R@1 23.6± 0.6 #8 of 11 Archive leaderboard report
Text to Audio Retrieval AudioCaps CE R@10 71.4±0.5 #8 of 11 Archive leaderboard report
Text to Audio Retrieval AudioCaps MoEE R@1 23.0±0.7 #10 of 11 Archive leaderboard report
Text to Audio Retrieval AudioCaps MoEE R@10 71.0±1.2 #10 of 11 Archive leaderboard report
Text to Audio Retrieval Clotho MMT R@1 6.5±0.6 #10 of 12 Archive leaderboard report
Text to Audio Retrieval Clotho MMT R@10 32.8±2.1 #10 of 12 Archive leaderboard report
Text to Audio Retrieval Clotho CE(pretraining:SoundDescs) R@1 6.4±0.5 #11 of 12 Archive leaderboard report
Text to Audio Retrieval Clotho CE(pretraining:SoundDescs) R@10 32.5±1.7 #11 of 12 Archive leaderboard report
Text to Audio Retrieval SoundDescs CE R@1 31.1±0.2 #1 of 4 Archive leaderboard report
Text to Audio Retrieval SoundDescs CE R@10 70.8±0.5 #1 of 4 Archive leaderboard report
Text to Audio Retrieval SoundDescs MoEE R@1 30.8±0.7 #2 of 4 Archive leaderboard report
Text to Audio Retrieval SoundDescs MoEE R@10 70.9±0.5 #2 of 4 Archive leaderboard report
Text to Audio Retrieval SoundDescs MMT R@1 30.7±0.4 #3 of 4 Archive leaderboard report
Text to Audio Retrieval SoundDescs MMT R@10 72.7±0.8 #3 of 4 Archive leaderboard report
Text to Audio Retrieval SoundDescs CE(pretrained: AudioCaps) R@1 23.3±0.7 #4 of 4 Archive leaderboard report
Text to Audio Retrieval SoundDescs CE(pretrained: AudioCaps) R@10 63.9±0.5 #4 of 4 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections