{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/2408-02657","title":"Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining","arxiv_id":"2408.02657","date":"2024-08-05","proceeding":null,"authors":["Dongyang Liu","Shitian Zhao","Le Zhuo","Weifeng Lin","Yu Qiao","Hongsheng Li","Peng Gao"],"abstract":"We present Lumina-mGPT, a family of multimodal autoregressive models capable of various vision and language tasks, particularly excelling in generating flexible photorealistic images from text descriptions. By initializing from multimodal Generative PreTraining (mGPT), we demonstrate that decoder-only Autoregressive (AR) model can achieve image generation performance comparable to modern diffusion models with high efficiency through Flexible Progressive Supervised Fine-tuning (FP-SFT). Equipped with our proposed Unambiguous image Representation (UniRep), Lumina-mGPT can flexibly generate high-quality images of varying aspect ratios. Building on the strong image generation capabilities, we further explore Ominiponent Supervised Fine-tuning (Omni-SFT), an initial attempt to elevate Lumina-mGPT into a unified multi-modal generalist. The resulting model demonstrates versatile multimodal capabilities, including visual generation tasks like text-to-image/multiview generation and controllable generation, visual recognition tasks like segmentation and depth estimation, and vision-language tasks like multi-turn visual question answering, showing the rosy potential of the technical direction. Codes and checkpoints are available at https://github.com/Alpha-VLLM/Lumina-mGPT.","url_abs":"https://arxiv.org/abs/2408.02657v2","url_pdf":"https://arxiv.org/pdf/2408.02657v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"2408-02657","repo_url":"https://github.com/alpha-vllm/lumina-mgpt","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}},{"paper_slug":"2408-02657","repo_url":"https://github.com/alpha-vllm/lumina-t2x","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"depth-estimation","task_name":"Depth Estimation"},{"task_slug":"image-generation","task_name":"Image Generation"},{"task_slug":"question-answering","task_name":"Question Answering"},{"task_slug":"text-to-image-generation-1","task_name":"Text to Image Generation"},{"task_slug":"text-to-image-generation","task_name":"Text-to-Image Generation"},{"task_slug":"visual-question-answering-1","task_name":"Visual Question Answering"}],"methods":[{"method_slug":"diffusion","method_name":"Diffusion"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2408.02657","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2408.02657"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/alpha-vllm/lumina-mgpt","reach":{"status":"ok","spdx":"Apache-2.0"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/alpha-vllm/lumina-t2x","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran_fixture":1,"ran_draft_wrong":4,"ran":2,"unverified":1},"by_repo_kind":{"official":{"samples":6,"ran":6,"repositories":1},"listed":{"samples":1,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":1,"samples":[{"code_sha256_prefix":"30d7eec482ebf6b1","entry":"repeat_kv","repo":"alpha-vllm/lumina-mgpt","repo_kind":"official","path":"lumina_mgpt/model/chameleon/modeling_chameleon.py","file_url":"https://github.com/alpha-vllm/lumina-mgpt/blob/HEAD/lumina_mgpt/model/chameleon/modeling_chameleon.py","link_basis":"harvester_set","language":"python","status":"ran_fixture","verification_level":2,"contract_check":"RAISES","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"30d7eec482ebf6b1"}},{"code_sha256_prefix":"2a92f0d95dce3ea2","entry":"add_options","repo":"alpha-vllm/lumina-t2x","repo_kind":"listed","path":"lumina_next_t2i/entry_point.py","file_url":"https://github.com/alpha-vllm/lumina-t2x/blob/HEAD/lumina_next_t2i/entry_point.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"2a92f0d95dce3ea2"}},{"code_sha256_prefix":"bac65c3dafaec040","entry":"apply_rotary_pos_emb","repo":"alpha-vllm/lumina-mgpt","repo_kind":"official","path":"lumina_mgpt/model/chameleon/modeling_chameleon.py","file_url":"https://github.com/alpha-vllm/lumina-mgpt/blob/HEAD/lumina_mgpt/model/chameleon/modeling_chameleon.py","link_basis":"harvester_set","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":"invariant","behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"bac65c3dafaec040"}},{"code_sha256_prefix":"507f475feb734647","entry":"compute_intermediate_size","repo":"alpha-vllm/lumina-mgpt","repo_kind":"official","path":"lumina_mgpt/model/chameleon/convert_chameleon_weights_to_hf.py","file_url":"https://github.com/alpha-vllm/lumina-mgpt/blob/HEAD/lumina_mgpt/model/chameleon/convert_chameleon_weights_to_hf.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"507f475feb734647"}},{"code_sha256_prefix":"6a0f498c92fe4fa3","entry":"make_batched_images","repo":"alpha-vllm/lumina-mgpt","repo_kind":"official","path":"lumina_mgpt/model/chameleon/image_processing_chameleon.py","file_url":"https://github.com/alpha-vllm/lumina-mgpt/blob/HEAD/lumina_mgpt/model/chameleon/image_processing_chameleon.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"6a0f498c92fe4fa3"}},{"code_sha256_prefix":"c5bcf01d18bba63d","entry":"read_json","repo":"alpha-vllm/lumina-mgpt","repo_kind":"official","path":"lumina_mgpt/model/chameleon/convert_chameleon_weights_to_hf.py","file_url":"https://github.com/alpha-vllm/lumina-mgpt/blob/HEAD/lumina_mgpt/model/chameleon/convert_chameleon_weights_to_hf.py","link_basis":"harvester_set","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"c5bcf01d18bba63d"}},{"code_sha256_prefix":"b99eea6376d1e212","entry":"rotate_half","repo":"alpha-vllm/lumina-mgpt","repo_kind":"official","path":"lumina_mgpt/model/chameleon/modeling_chameleon.py","file_url":"https://github.com/alpha-vllm/lumina-mgpt/blob/HEAD/lumina_mgpt/model/chameleon/modeling_chameleon.py","link_basis":"harvester_set","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"b99eea6376d1e212"}},{"code_sha256_prefix":"2fc6fdb85f13dc41","entry":"none_or_str","repo":null,"repo_kind":null,"path":null,"file_url":null,"link_basis":"identical_code_first_harvested_elsewhere","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":null,"inline_ok":false,"mcp_get_code":{"code_sha256":"2fc6fdb85f13dc41"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}