Browse State-of-the-Art › Audio Generation
Audio Generation
124 papers with code · 3 benchmarks · 11 datasets archive 2025-07-28
Audio generation (synthesis) is the task of generating raw audio such as speech.
( Image credit: MelNet )
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
3 leaderboard tables shown for this task, 3 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| AudioCaps (23 rows) | ETTA-FT-AC-100k | ETTA: Elucidating the Design Space of Text-to-Audio Models | code | — | Compare |
| Classical music, 5 seconds at 12 kHz (2 rows) | Sparse Transformer 152M (strided) | Generating Long Sequences with Sparse Transformers | code | Syntology ran 5 of 6 samples · 1 unverified | Compare |
| Symphony music (1 row) | SymphonyNet | Symphony Generation with Permutation Invariant Language Model | code | Syntology ran 2 of 2 samples · 0 unverified | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
11 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
4 subtasks in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 124 papers with code (270 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
12 Sep 2016 62 repositories listed Syntology ran 41 of 103 samples · 62 unverified · 25 pointer-only (licence)This paper introduces WaveNet, a deep neural network for generating raw audio waveforms.
-
12 Feb 2018 22 repositories listed Syntology ran 2 of 4 samples · 2 unverified · 4 pointer-only (licence)Audio signals are sampled at high temporal resolutions, and learning to synthesize audio requires capturing structure across a range of timescales.
-
7 Sep 2022 6 repositories listed Syntology ran 4 of 15 samples · 11 unverified · 3 pointer-only (licence)We introduce AudioLM, a framework for high-quality audio generation with long-term consistency.
-
20 Feb 2022 6 repositories listed Syntology ran 6 of 6 samples · 0 unverified · 4 pointer-only (licence)SaShiMi yields state-of-the-art performance for unconditional waveform generation in the autoregressive setting.
-
23 Feb 2019 6 repositories listedEfficient audio synthesis is an inherently difficult machine learning task, as human perception is sensitive to both global structure and fine-scale waveform coherence.
-
9 Jun 2022 5 repositories listed Syntology ran 8 of 17 samples · 9 unverified · 1 pointer-only (licence)Despite recent progress in generative adversarial network (GAN)-based vocoders, where the model generates raw waveform conditioned on acoustic features, it is challenging to synthesize high-fidelity audio for numerous…
-
4 Jun 2019 5 repositories listed Syntology ran 4 of 4 samples · 0 unverified · 4 pointer-only (licence)Capturing high-level structure in audio waveforms is challenging because a single second of audio spans tens of thousands of timesteps.
-
11 Jun 2023 4 repositories listed Syntology ran 27 of 38 samples · 11 unverifiedLanguage models have been successfully used to model natural signals, such as images, speech, and music.
-
29 Jan 2023 4 repositories listed Syntology ran 8 of 21 samples · 13 unverified · 10 pointer-only (licence)By learning the latent representations of audio signals and their compositions without modeling the cross-modal relationship, AudioLDM is advantageous in both generation quality and computational efficiency.
-
2 Aug 2017 4 repositories listed Syntology ran 4 of 4 samples · 0 unverified · 3 pointer-only (licence)We introduce a new audio processing technique that increases the sampling rate of signals such as speech or music using deep convolutional neural networks.
-
22 Dec 2016 4 repositories listed Syntology ran 0 of 11 samples · 11 unverifiedIn this paper we propose a novel model for unconditional audio generation based on generating one audio sample at a time.
-
16 May 2023 3 repositories listed Syntology ran 6 of 11 samples · 5 unverified · 3 pointer-only (licence)We present SoundStorm, a model for efficient, non-autoregressive audio generation.
-
18 Apr 2023 3 repositories listedKeywords: Bark, ai voice cloning, Suno, text-to-speech, artificial intelligence, audio generation, Meta's encodec, audio codebooks, semantic tokens, HuBert, transformer-based model, multilingual speech, wav2vec, linear…
-
18 Apr 2022 3 repositories listedJoint time-frequency scattering (JTFS) is a convolutional operator in the time-frequency domain which extracts spectrotemporal modulations at various rates and scales.
-
24 Mar 2022 3 repositories listed Syntology ran 9 of 12 samples · 3 unverified · 1 pointer-only (licence)Generative adversarial networks have recently demonstrated outstanding performance in neural vocoding outperforming best autoregressive and flow-based models.
-
17 Oct 2021 3 repositories listed Syntology ran 1 of 7 samples · 6 unverifiedIn this work, we propose a single model capable of generating visually relevant, high-fidelity sounds prompted with a set of frames from open-domain videos in less time than it takes to play it on a single GPU.
-
14 Jan 2020 3 repositories listed Syntology ran 0 of 5 samples · 5 unverified · 5 pointer-only (licence)In this paper, we introduce the Differentiable Digital Signal Processing (DDSP) library, which enables direct integration of classic signal processing elements with deep learning methods.
-
3 Jun 2019 3 repositories listed Syntology ran 1 of 5 samples · 4 unverifiedEnd-to-end models for raw audio generation are a challenge, specially if they have to work with non-parallel data, which is a desirable setup in many situations.
-
12 Apr 2019 3 repositories listedIts training data subsets can directly be visualized in the 3D latent representation.
-
7 Feb 2025 2 repositories listedTo address this issue, we propose Self-Loop Latent Swap, a frame-level bidirectional swap applied to the overlapping region of adjacent views.
-
28 Oct 2024 2 repositories listedTo enhance the sensitivity of deepfake audio features, we propose a deepfake audio detection model that incorporates an SLS (Sensitive Layer Selection) module.
-
17 Oct 2024 2 repositories listedOur models set a new state-of-the-art on multiple tasks: text-to-video synthesis, video personalization, video editing, video-to-audio generation, and text-to-audio generation.
-
24 Sep 2024 2 repositories listedNeural codecs have become crucial to recent speech and audio generation research.
-
1 Jun 2024 2 repositories listedTo overcome the convergence issue inherent in LDMs with reduced sample iterations, we propose the Guided Latent Consistency Distillation with a multi-step Ordinary Differential Equation (ODE) solver.
-
7 Feb 2024 2 repositories listed Syntology ran 2 of 4 samples · 2 unverifiedGenerating long-form 44.
-
22 Sep 2023 2 repositories listedDiffusion models have gained prominence in the image domain for their capabilities in data generation and transformation, achieving state-of-the-art performance in various tasks in both image and audio domains.
-
10 Aug 2023 2 repositories listed Syntology ran 12 of 27 samples · 15 unverified · 13 pointer-only (licence)Any audio can be translated into LOA based on AudioMAE, a self-supervised pre-trained representation learning model.
-
31 May 2023 2 repositories listedIn contrast, symbolic music offers ease of editing, making it more accessible for users to manipulate specific musical elements.
-
19 May 2023 2 repositories listed Syntology ran 2 of 13 samples · 11 unverifiedWe present Composable Diffusion (CoDi), a novel generative model capable of generating any combination of output modalities, such as language, image, video, or audio, from any combination of input modalities.
-
20 Dec 2021 2 repositories listed Syntology ran 0 of 2 samples · 2 unverifiedHigh-fidelity multi-singer singing voice synthesis is challenging for neural vocoder due to the singing voice data shortage, limited singer generalization, and large computational cost.
Syntology lines on 19 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections