Papers › Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

22 Aug 2024arXiv:2408.12528archive 2025-07-28

Jinheng Xie, Weijia Mao, Zechen Bai, David Junhao Zhang, WeiHao Wang, Kevin Qinghong Lin, YuChao Gu, Zhijie Chen, Zhenheng Yang, Mike Zheng Shou

We present a unified transformer, i.e., Show-o, that unifies multimodal understanding and generation. Unlike fully autoregressive models, Show-o unifies autoregressive and (discrete) diffusion modeling to adaptively handle inputs and outputs of various and mixed modalities. The unified model flexibly supports a wide range of vision-language tasks including visual question-answering, text-to-image generation, text-guided inpainting/extrapolation, and mixed-modality generation. Across various benchmarks, it demonstrates comparable or superior performance to existing individual models with an equivalent or larger number of parameters tailored for understanding or generation. This significantly highlights its potential as a next-generation foundation model. Code and models are released at https://github.com/showlab/Show-o.

PaperPDFCode

In Syntology View this paper on Syntology: its repositories, every harvested function with whether it ran, its licence and the call to fetch it.

Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

showlab/show-o mentioned in papermentioned on GitHubjaxApache-2.0 report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

10-shot image generationImage GenerationQuestion AnsweringText to Image GenerationText-to-Image GenerationVisual Question Answering

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
WISE Show-o Biology 0.30 #9 of 11 Archive leaderboard report
WISE Show-o Chemistry 0.30 #9 of 11 Archive leaderboard report
WISE Show-o Cultural 0.28 #9 of 11 Archive leaderboard report
WISE Show-o Overall 0.35 #9 of 11 Archive leaderboard report
WISE Show-o Physics 0.46 #9 of 11 Archive leaderboard report
WISE Show-o Space 0.48 #9 of 11 Archive leaderboard report
WISE Show-o Time 0.40 #9 of 11 Archive leaderboard report
Image Generation WISE Show-o Biology 0.30 #13 of 14 Archive leaderboard report
Image Generation WISE Show-o Chemistry 0.30 #13 of 14 Archive leaderboard report
Image Generation WISE Show-o Cultural 0.28 #13 of 14 Archive leaderboard report
Image Generation WISE Show-o Overall 0.35 #13 of 14 Archive leaderboard report
Image Generation WISE Show-o Physics 0.46 #13 of 14 Archive leaderboard report
Image Generation WISE Show-o Space 0.48 #13 of 14 Archive leaderboard report
Image Generation WISE Show-o Time 0.40 #13 of 14 Archive leaderboard report
Text-to-Image Generation GenEval Und. and Gen. Show-o (Ours) Overall 0.68 #14 of 20 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Diffusion

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections