Papers › ETTA: Elucidating the Design Space of Text-to-Audio Models

ETTA: Elucidating the Design Space of Text-to-Audio Models

26 Dec 2024arXiv:2412.19351archive 2025-07-28

Sang-gil Lee, Zhifeng Kong, Arushi Goel, Sungwon Kim, Rafael Valle, Bryan Catanzaro

Recent years have seen significant progress in Text-To-Audio (TTA) synthesis, enabling users to enrich their creative workflows with synthetic audio generated from natural language prompts. Despite this progress, the effects of data, model architecture, training objective functions, and sampling strategies on target benchmarks are not well understood. With the purpose of providing a holistic understanding of the design space of TTA models, we set up a large-scale empirical experiment focused on diffusion and flow matching models. Our contributions include: 1) AF-Synthetic, a large dataset of high quality synthetic captions obtained from an audio understanding model; 2) a systematic comparison of different architectural, training, and inference design choices for TTA models; 3) an analysis of sampling methods and their Pareto curves with respect to generation quality and inference speed. We leverage the knowledge obtained from this extensive analysis to propose our best model dubbed Elucidated Text-To-Audio (ETTA). When evaluated on AudioCaps and MusicCaps, ETTA provides improvements over the baselines trained on publicly available data, while being competitive with models trained on proprietary data. Finally, we show ETTA's improved ability to generate creative audio following complex and imaginative captions -- a task that is more challenging than current benchmarks.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Audio GenerationAudio captioningLanguage ModellingMusic GenerationText-to-Music Generation

1 archive task tag without a task page not shown.

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Audio Generation AudioCaps ETTA-FT-AC-100k CLAP_LAION 0.60 #1 of 23 Archive leaderboard report
Audio Generation AudioCaps ETTA-FT-AC-100k CLAP_MS 0.43 #1 of 23 Archive leaderboard report
Audio Generation AudioCaps ETTA-FT-AC-100k FAD 2.03 #1 of 23 Archive leaderboard report
Audio Generation AudioCaps ETTA-FT-AC-100k FD 10.10 #1 of 23 Archive leaderboard report
Audio Generation AudioCaps ETTA-FT-AC-100k FD_openl3 61.79 #1 of 23 Archive leaderboard report
Audio Generation AudioCaps ETTA-FT-AC-100k IS 14.29 #1 of 23 Archive leaderboard report
Audio Generation AudioCaps ETTA-FT-AC-100k KL_passt 1.13 #1 of 23 Archive leaderboard report
Audio Generation AudioCaps ETTA CLAP_LAION 0.54 #5 of 23 Archive leaderboard report
Audio Generation AudioCaps ETTA CLAP_MS 0.43 #5 of 23 Archive leaderboard report
Audio Generation AudioCaps ETTA FAD 2.51 #5 of 23 Archive leaderboard report
Audio Generation AudioCaps ETTA FD 13.12 #5 of 23 Archive leaderboard report
Audio Generation AudioCaps ETTA FD_openl3 80.13 #5 of 23 Archive leaderboard report
Audio Generation AudioCaps ETTA IS 14.36 #5 of 23 Archive leaderboard report
Audio Generation AudioCaps ETTA KL_passt 1.22 #5 of 23 Archive leaderboard report
Text-to-Music Generation MusicCaps ETTA CLAP_LAION 0.51 #4 of 21 Archive leaderboard report
Text-to-Music Generation MusicCaps ETTA CLAP_MS 0.53 #4 of 21 Archive leaderboard report
Text-to-Music Generation MusicCaps ETTA FAD 1.91 #4 of 21 Archive leaderboard report
Text-to-Music Generation MusicCaps ETTA FD 10.06 #4 of 21 Archive leaderboard report
Text-to-Music Generation MusicCaps ETTA FD_openl3 92.18 #4 of 21 Archive leaderboard report
Text-to-Music Generation MusicCaps ETTA IS 3.32 #4 of 21 Archive leaderboard report
Text-to-Music Generation MusicCaps ETTA KL_passt 0.84 #4 of 21 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

DiffusionSET

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections