Browse State-of-the-Art › Music Generation
Music Generation
190 papers with code · 1 benchmark · 33 datasets archive 2025-07-28
Musique guitar
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
1 leaderboard table shown for this task, 1 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| Song Describer Dataset (1 row) | OpenMusic | Quality-aware Masked Diffusion Transformer for Enhanced Music Generation | code | Syntology ran 4 of 7 samples · 3 unverified | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
33 datasets whose archive record lists this task, ordered by the archive's paper count. 30 shown of 33 until expanded.
Subtasks archive 2025-07-28
3 subtasks in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 190 papers with code (386 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
12 Sep 2018 12 repositories listed Syntology ran 2 of 4 samples · 2 unverified · 1 pointer-only (licence)This is impractical for long sequences such as musical compositions since their memory complexity for intermediate relative information is quadratic in the sequence length.
-
8 Jun 2023 8 repositories listed Syntology ran 11 of 18 samples · 7 unverified · 7 pointer-only (licence)We tackle the task of conditional music generation.
-
19 Sep 2017 8 repositories listedThe three models, which differ in the underlying assumptions and accordingly the network architectures, are referred to as the jamming model, the composer model and the hybrid model.
-
20 Feb 2022 6 repositories listed Syntology ran 6 of 6 samples · 0 unverified · 4 pointer-only (licence)SaShiMi yields state-of-the-art performance for unconditional waveform generation in the autoregressive setting.
-
26 Jan 2023 5 repositories listed Syntology ran 6 of 16 samples · 10 unverifiedWe introduce MusicLM, a model generating high-fidelity music from text descriptions such as "a calming violin melody backed by a distorted guitar riff".
-
9 Jun 2022 5 repositories listed Syntology ran 8 of 17 samples · 9 unverified · 1 pointer-only (licence)Despite recent progress in generative adversarial network (GAN)-based vocoders, where the model generates raw waveform conditioned on acoustic features, it is challenging to synthesize high-fidelity audio for numerous…
-
7 Jan 2021 5 repositories listed Syntology ran 3 of 9 samples · 6 unverifiedIn this paper, we present a conceptually different approach that explicitly takes into account the type of the tokens, such as note types and metric types.
-
4 Jun 2019 5 repositories listed Syntology ran 4 of 4 samples · 0 unverified · 4 pointer-only (licence)Capturing high-level structure in audio waveforms is challenging because a single second of audio spans tens of thousands of timesteps.
-
10 Aug 2018 5 repositories listedMusic generation has generally been focused on either creating scores or interpreting them.
-
14 Nov 2023 4 repositories listed Syntology ran 7 of 9 samples · 2 unverifiedThrough extensive experiments, we show that the quality of the music generated by Mustango is state-of-the-art, and the controllability through music-specific text prompts greatly outperforms other models such as…
-
18 Mar 2019 4 repositories listed Syntology ran 0 of 1 samples · 1 unverifiedMachine learning models of music typically break up the task of composition into a chronological process, composing a piece of music in a single pass from beginning to end.
-
29 Oct 2018 4 repositories listedGenerating musical audio directly with neural networks is notoriously difficult because it requires coherently modeling structure at many different timescales.
-
31 Mar 2017 4 repositories listedWe conduct a user study to compare the melody of eight-bar long generated by MidiNet and by Google's MelodyRNN models, each time using the same priming melody.
-
27 Jan 2023 3 repositories listedRecent years have seen the rapid development of large generative models for text; however, much less research has explored the connection between text and another "language" of communication -- music.
-
13 Aug 2020 3 repositories listed Syntology ran 0 of 9 samples · 9 unverifiedWe propose the Multi-Track Music Machine (MMM), a generative system based on the Transformer architecture that is capable of generating multi-track music.
-
25 Apr 2018 3 repositories listed Syntology ran 0 of 35 samples · 35 unverifiedExperimental results show that using binary neurons instead of HT or BS indeed leads to better results in a number of objective measures.
-
29 May 2015 3 repositories listedRecurrent neural networks (RNNs) are connectionist models that capture the dynamics of sequences via cycles in the network of nodes.
-
27 Mar 2025 2 repositories listedVision-to-music Generation, including video-to-music and image-to-music tasks, is a significant branch of multimodal artificial intelligence demonstrating vast application prospects in fields such as film scoring, short…
-
23 Sep 2024 2 repositories listedThe complex nature of musical emotion introduces inherent bias in both recognition and generation, particularly when relying on a single audio encoder, emotion classifier, or evaluation metric.
-
16 Sep 2024 2 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)We evaluate the proposed dataset by performing initial experiments regarding the detection and attribution of TTM-generated audio.
-
1 Sep 2024 2 repositories listed Syntology ran 6 of 9 samples · 3 unverified · 9 pointer-only (licence)This paper explores a simple extension of diffusion-based rectified flow Transformers for text-to-music generation, termed as FluxMusic.
-
21 Jul 2024 2 repositories listedHowever, textual prompts alone cannot precisely control temporal musical features such as chords and rhythm of the generated music.
-
24 May 2024 2 repositories listed Syntology ran 4 of 7 samples · 3 unverifiedText-to-music (TTM) generation, which converts textual descriptions into audio, opens up innovative avenues for multimedia creation.
-
16 May 2024 2 repositories listed Syntology ran 10 of 12 samples · 2 unverifiedA cascaded diffusion model is trained to model the hierarchical language, where each level is conditioned on its upper levels.
-
25 Apr 2024 2 repositories listed Syntology ran 2 of 5 samples · 3 unverified · 5 pointer-only (licence)We present COCOLA (Coherence-Oriented Contrastive Learning for Audio), a contrastive learning method for musical audio representations that captures the harmonic and rhythmic coherence between samples.
-
9 Aug 2023 2 repositories listedDespite the task's significance, prevailing generative models exhibit limitations in music quality, computational efficiency, and generalization.
-
31 May 2023 2 repositories listedIn contrast, symbolic music offers ease of editing, making it more accessible for users to manipulate specific musical elements.
-
27 Jan 2023 2 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)When used with deep learning, the symbolic music modality is often coupled with language model architectures.
-
21 Nov 2022 2 repositories listedBenefiting from large-scale datasets and pre-trained models, the field of generative models has recently gained significant momentum.
-
17 Nov 2022 2 repositories listedCommercial adoption of automatic music composition requires the capability of generating diverse and high-quality music suitable for the desired context (e.
Syntology lines on 17 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections