Papers › TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and...

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

30 Dec 2024arXiv:2412.21037archive 2025-07-28

Chia-Yu Hung, Navonil Majumder, Zhifeng Kong, Ambuj Mehrish, Amir Ali Bagherzadeh, Chuan Li, Rafael Valle, Bryan Catanzaro, Soujanya Poria

We introduce TangoFlux, an efficient Text-to-Audio (TTA) generative model with 515M parameters, capable of generating up to 30 seconds of 44.1kHz audio in just 3.7 seconds on a single A40 GPU. A key challenge in aligning TTA models lies in the difficulty of creating preference pairs, as TTA lacks structured mechanisms like verifiable rewards or gold-standard answers available for Large Language Models (LLMs). To address this, we propose CLAP-Ranked Preference Optimization (CRPO), a novel framework that iteratively generates and optimizes preference data to enhance TTA alignment. We demonstrate that the audio preference dataset generated using CRPO outperforms existing alternatives. With this framework, TangoFlux achieves state-of-the-art performance across both objective and subjective benchmarks. We open source all code and models to support further research in TTA generation.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

declare-lab/TangoFlux officialmentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Audio Generation

1 archive task tag without a task page not shown.

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Audio Generation AudioCaps TangoFlux CLAP_LAION 0.488 #2 of 23 Archive leaderboard report
Audio Generation AudioCaps TangoFlux FD_openl3 75.1 #2 of 23 Archive leaderboard report
Audio Generation AudioCaps TangoFlux IS 12.2 #2 of 23 Archive leaderboard report
Audio Generation AudioCaps TangoFlux KL_passt 1.15 #2 of 23 Archive leaderboard report
Audio Generation AudioCaps TangoFlux-base CLAP_LAION 0.438 #4 of 23 Archive leaderboard report
Audio Generation AudioCaps TangoFlux-base FD_openl3 79.7 #4 of 23 Archive leaderboard report
Audio Generation AudioCaps TangoFlux-base IS 10.7 #4 of 23 Archive leaderboard report
Audio Generation AudioCaps TangoFlux-base KL_passt 1.23 #4 of 23 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections