Browse State-of-the-Art › Resynthesis
Resynthesis
18 papers with code · 0 benchmarks · 2 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
2 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
18 shown of 18 papers with code (51 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
18 Apr 2022 3 repositories listedJoint time-frequency scattering (JTFS) is a convolutional operator in the time-frequency domain which extracts spectrotemporal modulations at various rates and scales.
-
16 Apr 2019 3 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)In this paper, we tackle the more realistic scenario where unexpected objects of unknown classes can appear at test time.
-
1 Apr 2021 2 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)We propose using self-supervised discrete representations for the task of speech resynthesis.
-
1 Feb 2021 2 repositories listedWe introduce Generative Spoken Language Modeling, the task of learning the acoustic and linguistic characteristics of a language from raw audio (no text, no labels), and a set of metrics to automatically evaluate the…
-
29 May 2025 1 repository listedWe investigate SLM performance as we vary codebook size and unit coarseness using the simple duration-penalized dynamic programming (DPDP) method.
-
9 Jan 2025 1 repository listedThis article introduces AnCoGen, a novel method that leverages a masked autoencoder to unify the analysis, control, and generation of speech signals within a single model.
-
21 Dec 2023 1 repository listedWe introduce EmphAssess, a prosodic benchmark designed to evaluate the capability of speech-to-speech models to encode and reproduce prosodic emphasis.
-
7 Jul 2023 1 repository listedUnsupervised object discovery (UOD) refers to the task of discriminating the whole region of objects from the background within a scene without relying on labeled datasets, which benefits the task of bounding-box-level…
-
14 Feb 2023 1 repository listedTo build speech processing methods that can handle speech as naturally as humans, researchers have explored multiple ways of building an invertible mapping from speech to an interpretable space.
-
2 Jan 2023 1 repository listedFollowing the findings of such an analysis, we propose practical improvements to the discrete unit for the GSLM.
-
24 Feb 2022 1 repository listedThis study focuses on the perception of music performances when contextual factors, such as room acoustics and instrument, change.
-
15 Feb 2022 1 repository listed Syntology ran 5 of 7 samples · 2 unverifiedTextless spoken language processing research aims to extend the applicability of standard NLP toolset onto spoken language and languages with few or no textual resources.
-
28 Aug 2020 1 repository listedRecently, a series of papers have presented different extensions of the VAE to process sequential data, which model not only the latent space but also the temporal dependencies within a sequence of data vectors and…
-
13 Aug 2020 1 repository listedThe presence of oscillations in aggregated COVID-19 data not only raises questions about the data's accuracy, it hinders understanding of the pandemic.
-
16 Jun 2019 1 repository listedWe propose to utilize the high quality speech generation capability of neural vocoders for noise suppression.
-
7 Mar 2019 1 repository listed Syntology ran 0 of 7 samples · 7 unverifiedIn this paper, we explore new approaches to combining information encoded within the learned representations of auto-encoders.
-
28 Nov 2018 1 repository listedSince the input photograph always observes only a part of the surface, we suggest a new inpainting method that completes the texture of the human body.
-
6 Nov 2018 1 repository listedIn audio signal processing, probabilistic time-frequency models have many benefits over their non-probabilistic counterparts.
Syntology lines on 4 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections