Browse State-of-the-Art › Music Captioning
Music Captioning
9 papers with code · 0 benchmarks · 5 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
5 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
9 shown of 9 papers with code (14 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
Music Understanding LLaMA: Advancing Text-to-Music Generation with Question Answering and Captioning22 Aug 2023 3 repositories listed Syntology ran 3 of 4 samples · 1 unverified · 4 pointer-only (licence)To fill this gap, we present a methodology for generating question-answer pairs from existing audio captioning datasets and introduce the MusicQA Dataset designed for answering open-ended music-related questions.
-
31 Jul 2023 2 repositories listed Syntology ran 4 of 7 samples · 3 unverified · 4 pointer-only (licence)In addition, we trained a transformer-based music captioning model with the dataset and evaluated it under zero-shot and transfer-learning settings.
-
18 Jun 2025 1 repository listedDetailed captions that accurately reflect the characteristics of a music piece can enrich music databases and drive forward research in music AI.
-
17 Sep 2024 1 repository listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)We evaluated the triplet-based musical knowledge for six general-purpose Transformer-based models.
-
29 Jul 2024 1 repository listedAugmented by the proposed synthetic dataset, FUTGA is enabled to identify the music's temporal changes at key transition points and their musical functions, as well as generate detailed descriptions for each music…
-
16 Nov 2023 1 repository listed Syntology ran 3 of 3 samples · 0 unverifiedWe introduce the Song Describer dataset (SDD), a new crowdsourced corpus of high-quality audio-caption pairs, designed for the evaluation of music-and-language models.
-
15 Sep 2023 1 repository listedLarge Language Models (LLMs) have shown immense potential in multimodal applications, yet the convergence of textual and musical domains remains not well-explored.
-
21 Dec 2022 1 repository listedMusic captioning has gained significant attention in the wake of the rising prominence of streaming media platforms.
-
24 Apr 2021 1 repository listedContent-based music information retrieval has seen rapid progress with the adoption of deep learning.
Syntology lines on 4 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections