{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/ranni-taming-text-to-image-diffusion-for","title":"Ranni: Taming Text-to-Image Diffusion for Accurate Instruction Following","arxiv_id":"2311.17002","date":"2023-11-28","proceeding":"CVPR 2024 1","authors":["Yutong Feng","Biao Gong","Di Chen","Yujun Shen","Yu Liu","Jingren Zhou"],"abstract":"Existing text-to-image (T2I) diffusion models usually struggle in interpreting complex prompts, especially those with quantity, object-attribute binding, and multi-subject descriptions. In this work, we introduce a semantic panel as the middleware in decoding texts to images, supporting the generator to better follow instructions. The panel is obtained through arranging the visual concepts parsed from the input text by the aid of large language models, and then injected into the denoising network as a detailed control signal to complement the text condition. To facilitate text-to-panel learning, we come up with a carefully designed semantic formatting protocol, accompanied by a fully-automatic data preparation pipeline. Thanks to such a design, our approach, which we call Ranni, manages to enhance a pre-trained T2I generator regarding its textual controllability. More importantly, the introduction of the generative middleware brings a more convenient form of interaction (i.e., directly adjusting the elements in the panel or using language instructions) and further allows users to finely customize their generation, based on which we develop a practical system and showcase its potential in continuous generation and chatting-based editing. Our project page is at https://ranni-t2i.github.io/Ranni.","url_abs":"https://arxiv.org/abs/2311.17002v3","url_pdf":"https://arxiv.org/pdf/2311.17002v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"ranni-taming-text-to-image-diffusion-for","repo_url":"https://github.com/lllyasviel/omost","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"attribute","task_name":"Attribute"},{"task_slug":"denoising","task_name":"Denoising"},{"task_slug":"instruction-following","task_name":"Instruction Following"}],"methods":[{"method_slug":"diffusion","method_name":"Diffusion"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2311.17002","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2311.17002"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/lllyasviel/omost","reach":{"status":"ok","spdx":"Apache-2.0"}}],"summary":{"unverified":7},"by_repo_kind":{"listed":{"samples":7,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"250a2dd0a13ccd39","entry":"binary_nonzero_positions","repo":"lllyasviel/omost","repo_kind":"listed","path":"lib_omost/canvas.py","file_url":"https://github.com/lllyasviel/omost/blob/HEAD/lib_omost/canvas.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"250a2dd0a13ccd39"}},{"code_sha256_prefix":"7856c3674a16dba8","entry":"closest_name","repo":"lllyasviel/omost","repo_kind":"listed","path":"lib_omost/canvas.py","file_url":"https://github.com/lllyasviel/omost/blob/HEAD/lib_omost/canvas.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"7856c3674a16dba8"}},{"code_sha256_prefix":"f006eb5d91bd2b9b","entry":"numpy2pytorch","repo":"lllyasviel/omost","repo_kind":"listed","path":"gradio_app.py","file_url":"https://github.com/lllyasviel/omost/blob/HEAD/gradio_app.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"f006eb5d91bd2b9b"}},{"code_sha256_prefix":"002306d269f88e44","entry":"pytorch2numpy","repo":"lllyasviel/omost","repo_kind":"listed","path":"gradio_app.py","file_url":"https://github.com/lllyasviel/omost/blob/HEAD/gradio_app.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"002306d269f88e44"}},{"code_sha256_prefix":"e16e386a3ad2a314","entry":"resize_without_crop","repo":"lllyasviel/omost","repo_kind":"listed","path":"gradio_app.py","file_url":"https://github.com/lllyasviel/omost/blob/HEAD/gradio_app.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"e16e386a3ad2a314"}},{"code_sha256_prefix":"68d038b2f7baeabf","entry":"safe_str","repo":"lllyasviel/omost","repo_kind":"listed","path":"lib_omost/canvas.py","file_url":"https://github.com/lllyasviel/omost/blob/HEAD/lib_omost/canvas.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"68d038b2f7baeabf"}},{"code_sha256_prefix":"eb5d84423fb8df2d","entry":"unload_all_models","repo":"lllyasviel/omost","repo_kind":"listed","path":"lib_omost/memory_management.py","file_url":"https://github.com/lllyasviel/omost/blob/HEAD/lib_omost/memory_management.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"eb5d84423fb8df2d"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}