{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/personalized-image-generation-with-large","title":"Personalized Image Generation with Large Multimodal Models","arxiv_id":"2410.14170","date":"2024-10-18","proceeding":null,"authors":["Yiyan Xu","Wenjie Wang","Yang Zhang","Biao Tang","Peng Yan","Fuli Feng","Xiangnan He"],"abstract":"Personalized content filtering, such as recommender systems, has become a critical infrastructure to alleviate information overload. However, these systems merely filter existing content and are constrained by its limited diversity, making it difficult to meet users' varied content needs. To address this limitation, personalized content generation has emerged as a promising direction with broad applications. Nevertheless, most existing research focuses on personalized text generation, with relatively little attention given to personalized image generation. The limited work in personalized image generation faces challenges in accurately capturing users' visual preferences and needs from noisy user-interacted images and complex multimodal instructions. Worse still, there is a lack of supervised data for training personalized image generation models. To overcome the challenges, we propose a Personalized Image Generation Framework named Pigeon, which adopts exceptional large multimodal models with three dedicated modules to capture users' visual preferences and needs from noisy user history and multimodal instructions. To alleviate the data scarcity, we introduce a two-stage preference alignment scheme, comprising masked preference reconstruction and pairwise preference alignment, to align Pigeon with the personalized image generation task. We apply Pigeon to personalized sticker and movie poster generation, where extensive quantitative results and human evaluation highlight its superiority over various generative baselines.","url_abs":"https://arxiv.org/abs/2410.14170v2","url_pdf":"https://arxiv.org/pdf/2410.14170v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"personalized-image-generation-with-large","repo_url":"https://github.com/yiyanxu/pigeon","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"image-generation","task_name":"Image Generation"},{"task_slug":"personalized-image-generation","task_name":"Personalized Image Generation"},{"task_slug":"recommendation-systems","task_name":"Recommendation Systems"},{"task_slug":"text-generation","task_name":"Text Generation"}],"methods":[{"method_slug":"align","method_name":"ALIGN"},{"method_slug":"attention","method_name":"Attention"},{"method_slug":"softmax","method_name":"Softmax"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2410.14170","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2410.14170"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/yiyanxu/pigeon","reach":{"status":"ok"}}],"summary":{"ran_draft_wrong":1,"ran":2,"unverified":2},"by_repo_kind":{"official":{"samples":5,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":5,"samples":[{"code_sha256_prefix":"dcd422d66b0581d8","entry":"get_cast_dtype","repo":"yiyanxu/pigeon","repo_kind":"official","path":"Pigeon/models/modeling_visual_encoder.py","file_url":"https://github.com/yiyanxu/pigeon/blob/HEAD/Pigeon/models/modeling_visual_encoder.py","link_basis":"harvester_set","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"dcd422d66b0581d8"}},{"code_sha256_prefix":"079b3d5ac6a75887","entry":"get_free_space","repo":"yiyanxu/pigeon","repo_kind":"official","path":"Pigeon/inference.py","file_url":"https://github.com/yiyanxu/pigeon/blob/HEAD/Pigeon/inference.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"079b3d5ac6a75887"}},{"code_sha256_prefix":"d9329c7762ffa525","entry":"split_tuple","repo":"yiyanxu/pigeon","repo_kind":"official","path":"Pigeon/models/lavit_for_pigeon.py","file_url":"https://github.com/yiyanxu/pigeon/blob/HEAD/Pigeon/models/lavit_for_pigeon.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"d9329c7762ffa525"}},{"code_sha256_prefix":"5651682f1a8791d1","entry":"build_eva_clip","repo":"yiyanxu/pigeon","repo_kind":"official","path":"Pigeon/models/modeling_visual_encoder.py","file_url":"https://github.com/yiyanxu/pigeon/blob/HEAD/Pigeon/models/modeling_visual_encoder.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"5651682f1a8791d1"}},{"code_sha256_prefix":"e98a8476a590a676","entry":"save_batch_results","repo":"yiyanxu/pigeon","repo_kind":"official","path":"Pigeon/inference.py","file_url":"https://github.com/yiyanxu/pigeon/blob/HEAD/Pigeon/inference.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"e98a8476a590a676"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}