Papers › Unite and Conquer: Plug & Play Multi-Modal Synthesis using Diffusion Models

Unite and Conquer: Plug & Play Multi-Modal Synthesis using Diffusion Models

1 Dec 2022CVPR 2023 1arXiv:2212.00793archive 2025-07-28

Nithin Gopalakrishnan Nair, Wele Gedara Chaminda Bandara, Vishal M. Patel

Generating photos satisfying multiple constraints find broad utility in the content creation industry. A key hurdle to accomplishing this task is the need for paired data consisting of all modalities (i.e., constraints) and their corresponding output. Moreover, existing methods need retraining using paired data across all modalities to introduce a new condition. This paper proposes a solution to this problem based on denoising diffusion probabilistic models (DDPMs). Our motivation for choosing diffusion models over other generative models comes from the flexible internal structure of diffusion models. Since each sampling step in the DDPM follows a Gaussian distribution, we show that there exists a closed-form solution for generating an image given various constraints. Our method can unite multiple diffusion models trained on multiple sub-tasks and conquer the combined task through our proposed sampling strategy. We also introduce a novel reliability parameter that allows using different off-the-shelf diffusion models trained across various datasets during sampling time alone to guide it to the desired outcome satisfying multiple constraints. We perform experiments on various standard multimodal tasks to demonstrate the effectiveness of our approach. More details can be found in https://nithin-gk.github.io/projectpages/Multidiff/index.html

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

Nithin-GK/UniteandConquer officialmentioned on GitHubpytorchApache-2.0 report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Face GenerationFace Sketch SynthesisImage GenerationSemantic SegmentationText-to-Face GenerationText-to-Image Generationmultimodal generation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Face Sketch Synthesis Multi-Modal CelebA-HQ Diffusion FID 26.09 #1 of 1 Archive leaderboard report
Text-to-Image Generation Multi-Modal-CelebA-HQ Unite and Conquer FID 26.09 #4 of 10 Archive leaderboard report
Text-to-Image Generation Multi-Modal-CelebA-HQ Unite and Conquer LPIPS 0.519 #4 of 10 Archive leaderboard report
multimodal generation Multi-Modal CelebA-HQ Diffusion FID 26.09 #1 of 1 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Diffusion

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections