{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/auffusion-leveraging-the-power-of-diffusion","title":"Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation","arxiv_id":"2401.01044","date":"2024-01-02","proceeding":null,"authors":["Jinlong Xue","Yayue Deng","Yingming Gao","Ya Li"],"abstract":"Recent advancements in diffusion models and large language models (LLMs) have significantly propelled the field of AIGC. Text-to-Audio (TTA), a burgeoning AIGC application designed to generate audio from natural language prompts, is attracting increasing attention. However, existing TTA studies often struggle with generation quality and text-audio alignment, especially for complex textual inputs. Drawing inspiration from state-of-the-art Text-to-Image (T2I) diffusion models, we introduce Auffusion, a TTA system adapting T2I model frameworks to TTA task, by effectively leveraging their inherent generative strengths and precise cross-modal alignment. Our objective and subjective evaluations demonstrate that Auffusion surpasses previous TTA approaches using limited data and computational resource. Furthermore, previous studies in T2I recognizes the significant impact of encoder choice on cross-modal alignment, like fine-grained details and object bindings, while similar evaluation is lacking in prior TTA works. Through comprehensive ablation studies and innovative cross-attention map visualizations, we provide insightful assessments of text-audio alignment in TTA. Our findings reveal Auffusion's superior capability in generating audios that accurately match textual descriptions, which further demonstrated in several related tasks, such as audio style transfer, inpainting and other manipulations. Our implementation and demos are available at https://auffusion.github.io.","url_abs":"https://arxiv.org/abs/2401.01044v1","url_pdf":"https://arxiv.org/pdf/2401.01044v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"auffusion-leveraging-the-power-of-diffusion","repo_url":"https://github.com/happylittlecat2333/Auffusion","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"audio-generation","task_name":"Audio Generation"},{"task_slug":"style-transfer","task_name":"Style Transfer"},{"task_slug":"cross-modal-alignment","task_name":"cross-modal alignment"}],"methods":[{"method_slug":"diffusion","method_name":"Diffusion"},{"method_slug":"pixel-prediction","method_name":"Inpainting"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/audio-generation-on-audiocaps","task":"Audio Generation","dataset":"AudioCaps","model":"Auffusion","rank_in_archive_order":15,"of":23,"metrics":{"FAD":"1.63","FD":"21.99"},"uses_additional_data":false},{"leaderboard":"/sota/audio-generation-on-audiocaps","task":"Audio Generation","dataset":"AudioCaps","model":"Auffusion-Full","rank_in_archive_order":16,"of":23,"metrics":{"FAD":"1.76","FD":"23.08"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/2401.01044","atlas_url":"https://app.syntology.ai/?focus=2401.01044","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2401.01044"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/happylittlecat2333/Auffusion","reach":null}],"summary":{"ran_draft_wrong":2},"by_repo_kind":{"official":{"samples":1,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":2,"samples":[{"code_sha256_prefix":"ce890b65c7821d08","entry":"json_load","repo":"happylittlecat2333/Auffusion","repo_kind":"official","path":"auffusion_pipeline.py","file_url":"https://github.com/happylittlecat2333/Auffusion/blob/HEAD/auffusion_pipeline.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"ce890b65c7821d08"}},{"code_sha256_prefix":"bea2d776a332f2b0","entry":"rescale_noise_cfg","repo":null,"repo_kind":null,"path":null,"file_url":null,"link_basis":"identical_code_first_harvested_elsewhere","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":null,"inline_ok":false,"mcp_get_code":{"code_sha256":"bea2d776a332f2b0"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}