Papers › Taming Data and Transformers for Audio Generation
Taming Data and Transformers for Audio Generation
Moayed Haji-Ali, Willi Menapace, Aliaksandr Siarohin, Guha Balakrishnan, Vicente Ordonez
The scalability of ambient sound generators is hindered by data scarcity, insufficient caption quality, and limited scalability in model architecture. This work addresses these challenges by advancing both data and model scaling. First, we propose an efficient and scalable dataset collection pipeline tailored for ambient audio generation, resulting in AutoReCap-XL, the largest ambient audio-text dataset with over 47 million clips. To provide high-quality textual annotations, we propose AutoCap, a high-quality automatic audio captioning model. By adopting a Q-Former module and leveraging audio metadata, AutoCap substantially enhances caption quality, reaching a CIDEr score of $83.2$, a 3.2% improvement over previous captioning models. Finally, we propose GenAu, a scalable transformer-based audio generation architecture that we scale up to 1.25B parameters. We demonstrate its benefits from data scaling with synthetic captions as well as model size scaling. When compared to baseline audio generators trained at similar size and data scale, GenAu obtains significant improvements of 4.7% in FAD score, 11.1% in IS, and 13.5% in CLAP score. Our code, model checkpoints, and dataset are publicly available.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Audio Generation | AudioCaps | GenAu-Large | CLAP_MS | 0.668 | #11 of 23 | Archive leaderboard | report |
| Audio Generation | AudioCaps | GenAu-Large | FAD | 1.21 | #11 of 23 | Archive leaderboard | report |
| Audio Generation | AudioCaps | GenAu-Large | FD | 16.51 | #11 of 23 | Archive leaderboard | report |
| Audio captioning | AudioCaps | AutoCap | CIDEr | 0.832 | #5 of 18 | Archive leaderboard | report |
| Audio captioning | AudioCaps | AutoCap | METEOR | 0.253 | #5 of 18 | Archive leaderboard | report |
| Audio captioning | AudioCaps | AutoCap | ROUGE | 0.518 | #5 of 18 | Archive leaderboard | report |
| Audio captioning | AudioCaps | AutoCap | ROUGE-L | 0.518 | #5 of 18 | Archive leaderboard | report |
| Audio captioning | AudioCaps | AutoCap | SPICE | 0.182 | #5 of 18 | Archive leaderboard | report |
| Audio captioning | AudioCaps | AutoCap | SPIDEr | 0.507 | #5 of 18 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections