Papers › MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO

19 May 2025arXiv:2505.13031archive 2025-07-28

Yicheng Xiao, Lin Song, Yukang Chen, Yingmin Luo, Yuxin Chen, Yukang Gan, Wei Huang, Xiu Li, Xiaojuan Qi, Ying Shan

Recent text-to-image systems face limitations in handling multimodal inputs and complex reasoning tasks. We introduce MindOmni, a unified multimodal large language model that addresses these challenges by incorporating reasoning generation through reinforcement learning. MindOmni leverages a three-phase training strategy: i) design of a unified vision language model with a decoder-only diffusion module, ii) supervised fine-tuning with Chain-of-Thought (CoT) instruction data, and iii) our proposed Reasoning Generation Policy Optimization (RGPO) algorithm, utilizing multimodal feedback to effectively guide policy updates. Experimental results demonstrate that MindOmni outperforms existing models, achieving impressive performance on both understanding and generation benchmarks, meanwhile showcasing advanced fine-grained reasoning generation capabilities, especially with mathematical reasoning instruction. All codes will be made public at \href{https://github.com/EasonXiao-888/MindOmni}{https://github.com/EasonXiao-888/MindOmni}.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

easonxiao-888/mindomni officialmentioned in papermentioned on GitHubpytorchNOASSERTION report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

DecoderImage GenerationLanguage ModelingLanguage ModellingLarge Language ModelMathematical ReasoningMultimodal Large Language ModelText-to-Image Generation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Image Generation WISE MindOmni (w/ cot) Biology 0.76 #1 of 14 Archive leaderboard report
Image Generation WISE MindOmni (w/ cot) Chemistry 0.52 #1 of 14 Archive leaderboard report
Image Generation WISE MindOmni (w/ cot) Cultural 0.75 #1 of 14 Archive leaderboard report
Image Generation WISE MindOmni (w/ cot) Overall 0.71 #1 of 14 Archive leaderboard report
Image Generation WISE MindOmni (w/ cot) Physics 0.72 #1 of 14 Archive leaderboard report
Image Generation WISE MindOmni (w/ cot) Space 0.76 #1 of 14 Archive leaderboard report
Image Generation WISE MindOmni (w/ cot) Time 0.70 #1 of 14 Archive leaderboard report
Image Generation WISE MindOmni (w/o cot) Biology 0.36 #10 of 14 Archive leaderboard report
Image Generation WISE MindOmni (w/o cot) Chemistry 0.32 #10 of 14 Archive leaderboard report
Image Generation WISE MindOmni (w/o cot) Cultural 0.40 #10 of 14 Archive leaderboard report
Image Generation WISE MindOmni (w/o cot) Overall 0.43 #10 of 14 Archive leaderboard report
Image Generation WISE MindOmni (w/o cot) Physics 0.52 #10 of 14 Archive leaderboard report
Image Generation WISE MindOmni (w/o cot) Space 0.62 #10 of 14 Archive leaderboard report
Image Generation WISE MindOmni (w/o cot) Time 0.38 #10 of 14 Archive leaderboard report
Text-to-Image Generation GenEval MindOmni Color Attri. 0.71 #3 of 20 Archive leaderboard report
Text-to-Image Generation GenEval MindOmni Colors 0.90 #3 of 20 Archive leaderboard report
Text-to-Image Generation GenEval MindOmni Counting 0.71 #3 of 20 Archive leaderboard report
Text-to-Image Generation GenEval MindOmni Overall 0.83 #3 of 20 Archive leaderboard report
Text-to-Image Generation GenEval MindOmni Position 0.71 #3 of 20 Archive leaderboard report
Text-to-Image Generation GenEval MindOmni Single Obj. 0.99 #3 of 20 Archive leaderboard report
Text-to-Image Generation GenEval MindOmni Two Obj. 0.94 #3 of 20 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Diffusion

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections