Browse State-of-the-Art › Audio Synthesis
Audio Synthesis
55 papers with code · 0 benchmarks · 2 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
2 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 55 papers with code (127 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
4 Mar 2018 35 repositories listed Syntology ran 2 of 10 samples · 8 unverified · 1 pointer-only (licence)Our results indicate that a simple convolutional architecture outperforms canonical recurrent networks such as LSTMs across a diverse range of tasks and datasets, while demonstrating longer effective memory.
-
29 Mar 2017 30 repositories listed Syntology ran 7 of 25 samples · 18 unverified · 6 pointer-only (licence)A text-to-speech synthesis system typically consists of multiple stages, such as a text analysis frontend, an acoustic model and an audio synthesis module.
-
12 Feb 2018 22 repositories listed Syntology ran 2 of 4 samples · 2 unverified · 4 pointer-only (licence)Audio signals are sampled at high temporal resolutions, and learning to synthesize audio requires capturing structure across a range of timescales.
-
23 Feb 2018 16 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 1 pointer-only (licence)The small number of weights in a Sparse WaveRNN makes it possible to sample high-fidelity audio on a mobile CPU in real time.
-
21 Sep 2020 11 repositories listed Syntology ran 20 of 33 samples · 13 unverifiedIn this work, we propose DiffWave, a versatile diffusion probabilistic model for conditional and unconditional waveform generation.
-
11 Apr 2024 8 repositories listedInfinite impulse response filters are an essential building block of many time-varying audio systems, such as audio effects and synthesisers.
-
5 Apr 2017 8 repositories listed Syntology ran 2 of 3 samples · 1 unverified · 3 pointer-only (licence)Generative models in vision have seen rapid progress due to algorithmic improvements and the availability of high-quality image datasets.
-
23 Feb 2019 6 repositories listedEfficient audio synthesis is an inherently difficult machine learning task, as human perception is sensitive to both global structure and fine-scale waveform coherence.
-
9 Jun 2022 5 repositories listed Syntology ran 8 of 17 samples · 9 unverified · 1 pointer-only (licence)Despite recent progress in generative adversarial network (GAN)-based vocoders, where the model generates raw waveform conditioned on acoustic features, it is challenging to synthesize high-fidelity audio for numerous…
-
7 Jul 2021 5 repositories listed Syntology ran 10 of 19 samples · 9 unverified · 1 pointer-only (licence)In this paper, we propose Conditional Score-based Diffusion models for Imputation (CSDI), a novel time series imputation method that utilizes score-based diffusion models conditioned on observed data.
-
7 Jun 2024 4 repositories listedTraining the linear prediction (LP) operator end-to-end for audio synthesis in modern deep learning frameworks is slow due to its recursive formulation.
-
1 Jun 2023 4 repositories listed Syntology ran 5 of 16 samples · 11 unverifiedRecent advancements in neural vocoding are predominantly driven by Generative Adversarial Networks (GANs) operating in the time-domain.
-
9 Nov 2021 3 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 1 pointer-only (licence)By leveraging a multi-band decomposition of the raw waveform, we show that our model is the first able to generate 48kHz audio signals, while simultaneously running 20 times faster than real-time on a standard laptop…
-
9 Aug 2020 3 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedWhile recent neural sequence-to-sequence models have greatly improved the quality of speech synthesis, there has not been a system capable of fast training, fast inference and high-quality audio synthesis at the same…
-
14 Jan 2020 3 repositories listed Syntology ran 0 of 5 samples · 5 unverified · 5 pointer-only (licence)In this paper, we introduce the Differentiable Digital Signal Processing (DDSP) library, which enables direct integration of classic signal processing elements with deep learning methods.
-
25 Feb 2017 3 repositories listed Syntology ran 0 of 8 samples · 8 unverifiedWe present Deep Voice, a production-quality text-to-speech system constructed entirely from deep neural networks.
-
1 Jun 2024 2 repositories listedTo overcome the convergence issue inherent in LDMs with reduced sample iterations, we propose the Guided Latent Consistency Distillation with a multi-step Ordinary Differential Equation (ODE) solver.
-
4 Jun 2021 2 repositories listedAlthough recent works on neural vocoder have improved the quality of synthesized audio, there still exists a gap between generated and ground-truth audio in frequency space.
-
12 Mar 2021 2 repositories listedIn this paper, we present a real-time implementation of the DDSP library embedded in a virtual synthesizer as a plug-in that can be used in a Digital Audio Workstation.
-
31 Oct 2018 2 repositories listed Syntology ran 2 of 7 samples · 5 unverifiedIn this paper we propose WaveGlow: a flow-based network capable of generating high quality speech from mel-spectrograms.
-
9 May 2025 1 repository listedModal methods for simulating vibrations of strings, membranes, and plates are widely used in acoustics and physically informed audio synthesis.
-
15 Jan 2025 1 repository listedHere we introduce a renormalization group-based diffusion model that leverages multiscale nature of data distributions for realizing a high-quality data generation.
-
19 Dec 2024 1 repository listed Syntology ran 0 of 4 samples · 4 unverifiedWe propose to synthesize high-quality and synchronized audio, given video and optional text conditions, using a novel multimodal joint training framework MMAudio.
-
2 Dec 2024 1 repository listed Syntology ran 2 of 7 samples · 5 unverified · 7 pointer-only (licence)We introduce OmniFlow, a novel generative model designed for any-to-any generation tasks such as text-to-image, text-to-audio, and audio-to-image synthesis.
-
Where are we in audio deepfake detection? A systematic analysis over generative and detection models6 Oct 2024 1 repository listedThrough extensive experiments, (1) we reveal the limitations of existing detection methods and demonstrate that foundation models exhibit stronger generalization capabilities, likely due to their model size and the…
-
19 Sep 2024 1 repository listedModeling the natural contour of fundamental frequency (F0) plays a critical role in music audio synthesis.
-
15 Jul 2024 1 repository listed Syntology ran 2 of 2 samples · 0 unverifiedLatent diffusion models have shown promising results in audio generation, making notable advancements over traditional methods.
-
27 Jun 2024 1 repository listedThe scalability of ambient sound generators is hindered by data scarcity, insufficient caption quality, and limited scalability in model architecture.
-
4 Mar 2024 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedDenoising diffusion probabilistic models (DDPMs) are becoming the leading paradigm for generative models.
-
23 Jan 2024 1 repository listedWe introduce an open-source platform that comprises DiffMoog and an end-to-end sound matching framework.
Syntology lines on 18 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections