Browse State-of-the-Art › FAD
FAD
24 papers with code · 0 benchmarks · 0 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
No dataset record in the archive lists this task.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
24 shown of 24 papers with code (62 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
2 Nov 2023 5 repositories listed Syntology ran 5 of 10 samples · 5 unverified · 1 pointer-only (licence)The growing popularity of generative music models underlines the need for perceptually relevant, objective music quality metrics.
-
23 Sep 2024 2 repositories listedThe complex nature of musical emotion introduces inherent bias in both recognition and generation, particularly when relying on a single audio encoder, emotion classifier, or evaluation metric.
-
11 Jun 2025 1 repository listedThis paper presents a tutorial-style survey and implementation guide of BemaGANv2, an advanced GAN-based vocoder designed for high-fidelity and long-term audio generation.
-
25 Apr 2025 1 repository listedThis paper introduces Drum One-Shot Extraction, a task in which the goal is to extract drum one-shots that are present in the music mixture.
-
28 Mar 2025 1 repository listedCreating high-quality sound effects from videos and text prompts requires precise alignment between visual and audio domains, both semantically and temporally, along with step-by-step guidance for professional audio…
-
20 Mar 2025 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedIn this work, we rigorously study the design space of reference-based divergence metrics for evaluating TTM models through (1) designing four synthetic meta-evaluations to measure sensitivity to particular musical…
-
3 Mar 2025 1 repository listedWe propose FlowDec, a neural full-band audio codec for general audio sampled at 48 kHz that combines non-adversarial codec training with a stochastic postfilter based on a novel conditional flow matching method.
-
21 Feb 2025 1 repository listedAlthough being widely adopted for evaluating generated audio signals, the Fr\'echet Audio Distance (FAD) suffers from significant limitations, including reliance on Gaussian assumptions, sensitivity to sample size, and…
-
10 Dec 2024 1 repository listed Syntology ran 0 of 1 samples · 1 unverifiedIn this paper we introduce the Frechet Music Distance (FMD), a novel evaluation metric for generative symbolic music models, inspired by the Frechet Inception Distance (FID) in computer vision and Frechet Audio Distance…
-
10 Sep 2024 1 repository listed Syntology ran 6 of 6 samples · 0 unverifiedIts goal is to use one single diffusion model to generate mutually-coherent music sources, that are then mixed to form the music.
-
16 Aug 2024 1 repository listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)Audio generation has achieved remarkable progress with the advance of sophisticated generative models, such as diffusion models (DMs) and autoregressive (AR) models.
-
7 Aug 2024 1 repository listed Syntology ran 3 of 8 samples · 5 unverifiedHowever, the fusion of LiDAR and 4D radar is challenging because they differ significantly in terms of data quality and the degree of degradation in adverse weather.
-
27 Jun 2024 1 repository listedThe scalability of ambient sound generators is hindered by data scarcity, insufficient caption quality, and limited scalability in model architecture.
-
7 Jun 2024 1 repository listed Syntology ran 12 of 12 samples · 0 unverified · 12 pointer-only (licence)MeLFusion is a text-to-music diffusion model with a novel "visual synapse", which effectively infuses the semantics from the visual modality into the generated music.
-
29 May 2024 1 repository listedThis research contributes to the comprehension of the human auditory system, pushing boundaries in neural decoding and audio reconstruction methodologies.
-
18 Mar 2024 1 repository listedWe introduce a new loss term to enhance Foley sound generation in AudioLDM without post-filtering.
-
23 Aug 2023 1 repository listed Syntology ran 2 of 6 samples · 4 unverified · 6 pointer-only (licence)In this paper, we present a novel Amplitude-Modulated Stochastic Perturbation and Vortex Convolutional Network, AMSP-UOD, designed for underwater object detection.
-
14 Mar 2023 1 repository listedA popular approach to streaming speech translation is to employ a single offline model with a wait-k policy to support different latency requirements, which is simpler than training multiple online models with different…
-
19 Dec 2022 1 repository listed Syntology ran 9 of 18 samples · 9 unverifiedTo generate joint audio-video pairs, we propose a novel Multi-Modal Diffusion model (i.
-
28 Nov 2022 1 repository listedIn this paper, we introduce a novel Refined Semantic enhancement method towards Frequency Diffusion (RSFD), a captioning model that constantly perceives the linguistic representation of the infrequent tokens.
-
25 Jun 2022 1 repository listedWe describe our approach for the generative emotional vocal burst task (ExVo Generate) of the ICML Expressive Vocalizations Competition.
-
5 Sep 2021 1 repository listedThis research project investigates the application of deep learning to timbre transfer, where the timbre of a source audio can be converted to the timbre of a target audio with minimal loss in quality.
-
23 Jul 2020 1 repository listedFAD consists of a designed search space and an efficient architecture search algorithm.
-
17 Feb 2019 1 repository listedWith the increasing importance of online communities, discussion forums, and customer reviews, Internet "trolls" have proliferated thereby making it difficult for information seekers to find relevant and correct…
Syntology lines on 9 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections