{"url":"/task/multimodal-generation","name":"multimodal generation","slug":"multimodal-generation","description_markdown":"**Multimodal generation** refers to the process of generating outputs that incorporate multiple modalities, such as images, text, and sound. This can be done using deep learning models that are trained on data that includes multiple modalities, allowing the models to generate output that is informed by more than one type of data.\r\n\r\nFor example, a multimodal generation model could be trained to generate captions for images that incorporate both text and visual information. The model could learn to identify objects in the image and generate descriptions of them in natural language, while also taking into account contextual information and the relationships between the objects in the image.\r\n\r\nMultimodal generation can also be used in other applications, such as generating realistic images from textual descriptions or generating audio descriptions of video content. By combining multiple modalities in this way, multimodal generation models can produce more accurate and comprehensive output, making them useful for a wide range of applications.","categories":[{"name":"Natural Language Processing","url":"/area/natural-language-processing"},{"name":"Time Series","url":"/area/time-series"}],"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","slug_source":"archive_url"},"counts":{"papers_tagged":98,"papers_with_code":53,"benchmarks":1,"benchmark_tables_in_archive":1,"benchmark_tables_shown":1,"benchmark_tables_withheld_as_spam":0,"benchmark_definition":"a leaderboard table with at least one row; benchmark_tables_shown also counts the zero-row tables; benchmark_tables_in_archive adds the tables withheld as spam","datasets":6,"subtasks":0,"parent_tasks":1},"benchmarks":[{"leaderboard":"/sota/multimodal-generation-on-multi-modal-celeba","slug":"multimodal-generation-on-multi-modal-celeba","dataset":"Multi-Modal CelebA-HQ","dataset_url":"/dataset/multi-modal-celeba-hq-1","rows_in_archive":1,"metrics":["FID"],"first_row_in_archive_order":{"model":"Diffusion","paper_title":"Unite and Conquer: Plug & Play Multi-Modal Synthesis using Diffusion Models","paper_url":"/paper/unite-and-conquer-cross-dataset-multimodal","paper_date":"2022-12-01","arxiv_id":"2212.00793","code_links":[{"title":"Nithin-GK/UniteandConquer","url":"https://github.com/Nithin-GK/UniteandConquer"}],"syntology":null}}],"datasets":[{"url":"/dataset/multi-modal-celeba-hq-1","name":"Multi-Modal CelebA-HQ","full_name":"","num_papers_in_archive":27},{"url":"/dataset/mmneedle","name":"MMNeedle","full_name":"Multimodal Needle in a Haystack","num_papers_in_archive":12},{"url":"/dataset/alex-20","name":"Alex-20","full_name":"","num_papers_in_archive":1},{"url":"/dataset/medtrinity-25m","name":"MedTrinity-25M","full_name":"","num_papers_in_archive":1},{"url":"/dataset/rebus","name":"REBUS","full_name":"A Robust Evaluation Benchmark of Understanding Symbols","num_papers_in_archive":1},{"url":"/dataset/taste-music-dataset","name":"taste-music-dataset","full_name":"Taste Music Dataset","num_papers_in_archive":1}],"subtasks":[],"parent_tasks":[{"url":"/task/multimodal-association","name":"Multimodal Association"}],"papers":{"order":"repositories listed in the archive (desc), then date (desc); the archive holds no stars","population":"papers tagged with this task that list at least one repository in the archive","shown":30,"of":53,"tagged_in_all":98,"items":[{"url":"/paper/pmg-personalized-multimodal-generation-with","title":"PMG : Personalized Multimodal Generation with Large Language Models","date":"2024-04-07","arxiv_id":"2404.08677","repositories_listed":3,"syntology":{"n":5,"n_ran":4,"n_unverified":1,"n_pointer_only":5}},{"url":"/paper/retrieval-augmented-generation-for-ai","title":"Retrieval-Augmented Generation for AI-Generated Content: A Survey","date":"2024-02-29","arxiv_id":"2402.19473","repositories_listed":3,"syntology":null},{"url":"/paper/finite-scalar-quantization-vq-vae-made-simple","title":"Finite Scalar Quantization: VQ-VAE Made Simple","date":"2023-09-27","arxiv_id":"2309.15505","repositories_listed":3,"syntology":{"n":3,"n_ran":2,"n_unverified":1,"n_pointer_only":2}},{"url":"/paper/emerging-properties-in-unified-multimodal","title":"Emerging Properties in Unified Multimodal Pretraining","date":"2025-05-20","arxiv_id":"2505.14683","repositories_listed":2,"syntology":{"n":21,"n_ran":9,"n_unverified":12,"n_pointer_only":0}},{"url":"/paper/vision-to-music-generation-a-survey","title":"Vision-to-Music Generation: A Survey","date":"2025-03-27","arxiv_id":"2503.21254","repositories_listed":2,"syntology":null},{"url":"/paper/new-benchmarks-for-accountable-text-based","title":"Accountable Textual-Visual Chat Learns to Reject Human Instructions in Image Re-creation","date":"2023-03-10","arxiv_id":"2303.05983","repositories_listed":2,"syntology":null},{"url":"/paper/grounding-language-models-to-images-for","title":"Grounding Language Models to Images for Multimodal Inputs and Outputs","date":"2023-01-31","arxiv_id":"2301.13823","repositories_listed":2,"syntology":{"n":8,"n_ran":2,"n_unverified":6,"n_pointer_only":0}},{"url":"/paper/gans-n-roses-stable-controllable-diverse","title":"GANs N' Roses: Stable, Controllable, Diverse Image to Image Translation (works for videos too!)","date":"2021-06-11","arxiv_id":"2106.06561","repositories_listed":2,"syntology":null},{"url":"/paper/omnigen2-exploration-to-advanced-multimodal","title":"OmniGen2: Exploration to Advanced Multimodal Generation","date":"2025-06-23","arxiv_id":"2506.18871","repositories_listed":1,"syntology":{"n":9,"n_ran":1,"n_unverified":8,"n_pointer_only":0}},{"url":"/paper/muddit-liberating-generation-beyond-text-to","title":"Muddit: Liberating Generation Beyond Text-to-Image with a Unified Discrete Diffusion Model","date":"2025-05-29","arxiv_id":"2505.23606","repositories_listed":1,"syntology":{"n":3,"n_ran":2,"n_unverified":1,"n_pointer_only":3}},{"url":"/paper/omnigenbench-a-benchmark-for-omnipotent","title":"OmniGenBench: A Benchmark for Omnipotent Multimodal Generation across 50+ Tasks","date":"2025-05-24","arxiv_id":"2505.18775","repositories_listed":1,"syntology":null},{"url":"/paper/an-empirical-study-of-gpt-4o-image-generation","title":"An Empirical Study of GPT-4o Image Generation Capabilities","date":"2025-04-08","arxiv_id":"2504.05979","repositories_listed":1,"syntology":null},{"url":"/paper/crystalformer-rl-reinforcement-fine-tuning","title":"CrystalFormer-RL: Reinforcement Fine-Tuning for Materials Design","date":"2025-04-03","arxiv_id":"2504.02367","repositories_listed":1,"syntology":{"n":3,"n_ran":2,"n_unverified":1,"n_pointer_only":2}},{"url":"/paper/fusdreamer-label-efficient-remote-sensing","title":"FusDreamer: Label-efficient Remote Sensing World Model for Multimodal Data Classification","date":"2025-03-18","arxiv_id":"2503.13814","repositories_listed":1,"syntology":null},{"url":"/paper/multimodal-chain-of-thought-reasoning-a","title":"Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey","date":"2025-03-16","arxiv_id":"2503.12605","repositories_listed":1,"syntology":null},{"url":"/paper/omnimamba-efficient-and-unified-multimodal","title":"OmniMamba: Efficient and Unified Multimodal Understanding and Generation via State Space Models","date":"2025-03-11","arxiv_id":"2503.08686","repositories_listed":1,"syntology":{"n":8,"n_ran":5,"n_unverified":3,"n_pointer_only":0}},{"url":"/paper/unified-reward-model-for-multimodal","title":"Unified Reward Model for Multimodal Understanding and Generation","date":"2025-03-07","arxiv_id":"2503.05236","repositories_listed":1,"syntology":null},{"url":"/paper/wegen-a-unified-model-for-interactive","title":"WeGen: A Unified Model for Interactive Multimodal Generation as We Chat","date":"2025-03-03","arxiv_id":"2503.01115","repositories_listed":1,"syntology":null},{"url":"/paper/ask-in-any-modality-a-comprehensive-survey-on","title":"Ask in Any Modality: A Comprehensive Survey on Multimodal Retrieval-Augmented Generation","date":"2025-02-12","arxiv_id":"2502.08826","repositories_listed":1,"syntology":null},{"url":"/paper/unicms-a-unified-consistency-model-for","title":"UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding","date":"2025-02-08","arxiv_id":"2502.05415","repositories_listed":1,"syntology":null},{"url":"/paper/mramg-bench-a-beyondtext-benchmark-for","title":"MRAMG-Bench: A Comprehensive Benchmark for Advancing Multimodal Retrieval-Augmented Multimodal Generation","date":"2025-02-06","arxiv_id":"2502.04176","repositories_listed":1,"syntology":null},{"url":"/paper/anid-how-far-are-we-evaluating-the","title":"D-Judge: How Far Are We? Evaluating the Discrepancies Between AI-synthesized Images and Natural Images through Multimodal Guidance","date":"2024-12-23","arxiv_id":"2412.17632","repositories_listed":1,"syntology":null},{"url":"/paper/multimodal-latent-language-modeling-with-next","title":"Multimodal Latent Language Modeling with Next-Token Diffusion","date":"2024-12-11","arxiv_id":"2412.08635","repositories_listed":1,"syntology":null},{"url":"/paper/gate-opening-a-comprehensive-benchmark-for","title":"OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation","date":"2024-11-27","arxiv_id":"2411.18499","repositories_listed":1,"syntology":null},{"url":"/paper/multi-modal-retrieval-augmented-multi-modal","title":"Multi-modal Retrieval Augmented Multi-modal Generation: A Benchmark, Evaluate Metrics and Strong Baselines","date":"2024-11-25","arxiv_id":"2411.16365","repositories_listed":1,"syntology":{"n":1,"n_ran":0,"n_unverified":1,"n_pointer_only":0}},{"url":"/paper/efficient-diffusion-models-a-comprehensive","title":"Efficient Diffusion Models: A Comprehensive Survey from Principles to Practices","date":"2024-10-15","arxiv_id":"2410.11795","repositories_listed":1,"syntology":null},{"url":"/paper/mm2latent-text-to-facial-image-generation-and","title":"MM2Latent: Text-to-facial image generation and editing in GANs with multimodal assistance","date":"2024-09-17","arxiv_id":"2409.11010","repositories_listed":1,"syntology":null},{"url":"/paper/pixelbytes-catching-unified-representation","title":"PixelBytes: Catching Unified Representation for Multimodal Generation","date":"2024-09-16","arxiv_id":"2410.01820","repositories_listed":1,"syntology":null},{"url":"/paper/pixelbytes-catching-unified-embedding-for","title":"PixelBytes: Catching Unified Embedding for Multimodal Generation","date":"2024-09-03","arxiv_id":"2409.15512","repositories_listed":1,"syntology":null},{"url":"/paper/unifashion-a-unified-vision-language-model","title":"UniFashion: A Unified Vision-Language Model for Multimodal Fashion Retrieval and Generation","date":"2024-08-21","arxiv_id":"2408.11305","repositories_listed":1,"syntology":{"n":11,"n_ran":8,"n_unverified":3,"n_pointer_only":11}}],"syntology_records":10,"syntology_note":"a paper without a record is not a recorded non-run: it may lack an arXiv id or simply be absent from the graph layer"},"description_links":{"kept":0,"unwrapped_to_text":0,"bare_urls_linked":0,"relative_images_dropped":0,"rule":"internal links are kept only when the target slug exists in the catalog"},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per-sample execution status on synthesized fixtures ('ran N of M samples'); not a correctness claim and not a ranking signal.","status_vocabulary":{"ran_honours":"ran, honoured the contract we drafted","ran_violates":"ran, violated the contract we drafted","ran_draft_wrong":"ran; our contract draft was wrong, not the code","ran_fixture":"ran; our fixture could not drive it","ran":"ran on a synthesized input","unverified":"unverified (harvested, no recorded run)"}},"not_shown":{"libraries":"the archive has no per-task library table","trend_sparklines":"the Trend column of the benchmarks table was a rendered image; it is not in the archive","social_and_latest_sorts":"stars and social signals are not in the archive"}}