Browse State-of-the-Art › multimodal generation
multimodal generation
53 papers with code · 1 benchmark · 6 datasets archive 2025-07-28
Multimodal generation refers to the process of generating outputs that incorporate multiple modalities, such as images, text, and sound. This can be done using deep learning models that are trained on data that includes multiple modalities, allowing the models to generate output that is informed by more than one type of data.
For example, a multimodal generation model could be trained to generate captions for images that incorporate both text and visual information. The model could learn to identify objects in the image and generate descriptions of them in natural language, while also taking into account contextual information and the relationships between the objects in the image.
Multimodal generation can also be used in other applications, such as generating realistic images from textual descriptions or generating audio descriptions of video content. By combining multiple modalities in this way, multimodal generation models can produce more accurate and comprehensive output, making them useful for a wide range of applications.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
1 leaderboard table shown for this task, 1 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| Multi-Modal CelebA-HQ (1 row) | Diffusion | Unite and Conquer: Plug & Play Multi-Modal Synthesis using Diffusion Models | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
6 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 53 papers with code (98 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
7 Apr 2024 3 repositories listed Syntology ran 4 of 5 samples · 1 unverified · 5 pointer-only (licence)Such user preferences are then fed into a generator, such as a multimodal LLM or diffusion model, to produce personalized content.
-
29 Feb 2024 3 repositories listedWe first classify RAG foundations according to how the retriever augments the generator, distilling the fundamental abstractions of the augmentation methodologies for various retrievers and generators.
-
27 Sep 2023 3 repositories listed Syntology ran 2 of 3 samples · 1 unverified · 2 pointer-only (licence)Each dimension is quantized to a small set of fixed values, leading to an (implicit) codebook given by the product of these sets.
-
20 May 2025 2 repositories listed Syntology ran 9 of 21 samples · 12 unverifiedUnifying multimodal understanding and generation has shown impressive capabilities in cutting-edge proprietary systems.
-
27 Mar 2025 2 repositories listedVision-to-music Generation, including video-to-music and image-to-music tasks, is a significant branch of multimodal artificial intelligence demonstrating vast application prospects in fields such as film scoring, short…
-
10 Mar 2023 2 repositories listedThis stream is subsequently fed into the decoder-based transformer to generate visual re-creations and textual feedback in the second stage.
-
31 Jan 2023 2 repositories listed Syntology ran 2 of 8 samples · 6 unverifiedWe propose an efficient method to ground pretrained text-only language models to the visual domain, enabling them to process arbitrarily interleaved image-and-text data, and generate text interleaved with retrieved…
-
11 Jun 2021 2 repositories listedThis adversarial loss guarantees the map is diverse -- a very wide range of anime can be produced from a single content code.
-
23 Jun 2025 1 repository listed Syntology ran 1 of 9 samples · 8 unverifiedTo facilitate the training of OmniGen2, we developed comprehensive data construction pipelines, encompassing image editing and in-context generation data.
-
29 May 2025 1 repository listed Syntology ran 2 of 3 samples · 1 unverified · 3 pointer-only (licence)Unlike prior unified diffusion models trained from scratch, Muddit integrates strong visual priors from a pretrained text-to-image backbone with a lightweight text decoder, enabling flexible and high-quality multimodal…
-
24 May 2025 1 repository listedRecent breakthroughs in large multimodal models (LMMs), such as the impressive GPT-4o-Native, have demonstrated remarkable proficiency in following general-purpose instructions for image generation.
-
8 Apr 2025 1 repository listedThe landscape of image generation has rapidly evolved, from early GAN-based approaches to diffusion models and, most recently, to unified generative architectures that seek to bridge understanding and generation tasks.
-
3 Apr 2025 1 repository listed Syntology ran 2 of 3 samples · 1 unverified · 2 pointer-only (licence)Reinforcement fine-tuning has instrumental enhanced the instruction-following and reasoning abilities of large language models.
-
18 Mar 2025 1 repository listedWorld models significantly enhance hierarchical understanding, improving data integration and learning efficiency.
-
16 Mar 2025 1 repository listedBy extending the advantage of chain-of-thought (CoT) reasoning in human-like step-by-step processes to multimodal contexts, multimodal CoT (MCoT) reasoning has recently garnered significant research attention,…
-
11 Mar 2025 1 repository listed Syntology ran 5 of 8 samples · 3 unverifiedThe model fully leverages Mamba-2's high computational and memory efficiency, extending its capabilities from text generation to multimodal generation.
-
7 Mar 2025 1 repository listedRecent advances in human preference alignment have significantly enhanced multimodal generation and understanding.
-
3 Mar 2025 1 repository listedIn this work, we introduce WeGen, a model that unifies multimodal generation and understanding, and promotes their interplay in iterative generation.
-
12 Feb 2025 1 repository listedLarge Language Models (LLMs) struggle with hallucinations and outdated knowledge due to their reliance on static training data.
-
8 Feb 2025 1 repository listedHowever, the key challenge is establishing a unified denoising perspective for both image and text generation, which is essential for establishing the consistency mapping.
-
6 Feb 2025 1 repository listedTo address this, we introduce the Multimodal Retrieval-Augmented Multimodal Generation (MRAMG) task, in which we aim to generate multimodal answers that combine both text and images, fully leveraging the multimodal data…
-
23 Dec 2024 1 repository listedIn Artificial Intelligence Generated Content (AIGC), distinguishing AI-synthesized images from natural ones remains a key challenge.
-
11 Dec 2024 1 repository listedIn this work, we propose Latent Language Modeling (LatentLM), which seamlessly integrates continuous and discrete data using causal Transformers.
-
27 Nov 2024 1 repository listedWhile the progress in unified models offers new solutions, existing benchmarks are insufficient for evaluating these methods due to limitations in data size and diversity.
-
25 Nov 2024 1 repository listed Syntology ran 0 of 1 samples · 1 unverifiedThis paper investigates an intriguing task of Multi-modal Retrieval Augmented Multi-modal Generation (M²RAG).
-
15 Oct 2024 1 repository listedAs one of the most popular and sought-after generative models in the recent years, diffusion models have sparked the interests of many researchers and steadily shown excellent advantage in various generative tasks such…
-
17 Sep 2024 1 repository listedIn this paper, we propose a practical framework - MM2Latent - for multimodal image generation and editing.
-
16 Sep 2024 1 repository listedThis report presents PixelBytes, an approach for unified multimodal representation learning.
-
3 Sep 2024 1 repository listedThis report introduces PixelBytes Embedding, a novel approach for unified multimodal representation learning.
-
21 Aug 2024 1 repository listed Syntology ran 8 of 11 samples · 3 unverified · 11 pointer-only (licence)In this work, we present UniFashion, a unified framework that simultaneously tackles the challenges of multimodal generation and retrieval tasks within the fashion domain, integrating image generation with retrieval…
Syntology lines on 10 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections