Papers › Differentiable Time-Varying Linear Prediction in the Context of End-to-End...

Differentiable Time-Varying Linear Prediction in the Context of End-to-End Analysis-by-Synthesis

7 Jun 2024arXiv:2406.05128archive 2025-07-28

Chin-Yun Yu, György Fazekas

Training the linear prediction (LP) operator end-to-end for audio synthesis in modern deep learning frameworks is slow due to its recursive formulation. In addition, frame-wise approximation as an acceleration method cannot generalise well to test time conditions where the LP is computed sample-wise. Efficient differentiable sample-wise LP for end-to-end training is the key to removing this barrier. We generalise the efficient time-invariant LP implementation from the GOLF vocoder to time-varying cases. Combining this with the classic source-filter model, we show that the improved GOLF learns LP coefficients and reconstructs the voice better than its frame-wise counterparts. Moreover, in our listening test, synthesised outputs from GOLF scored higher in quality ratings than the state-of-the-art differentiable WORLD vocoder.

PaperPDFCode

Code

yoyololicon/golf officialmentioned on GitHubpytorch report
yoyololicon/torchlpc officialmentioned on GitHubpytorch report
diffapf/torchlpc mentioned on GitHubpytorch report
iamycy/golf mentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Audio Synthesis

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections