Papers › Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

17 Oct 2024CVPR 2025 1arXiv:2410.13848archive 2025-07-28

Chengyue Wu, Xiaokang Chen, Zhiyu Wu, Yiyang Ma, Xingchao Liu, Zizheng Pan, Wen Liu, Zhenda Xie, Xingkai Yu, Chong Ruan, Ping Luo

In this paper, we introduce Janus, an autoregressive framework that unifies multimodal understanding and generation. Prior research often relies on a single visual encoder for both tasks, such as Chameleon. However, due to the differing levels of information granularity required by multimodal understanding and generation, this approach can lead to suboptimal performance, particularly in multimodal understanding. To address this issue, we decouple visual encoding into separate pathways, while still leveraging a single, unified transformer architecture for processing. The decoupling not only alleviates the conflict between the visual encoder's roles in understanding and generation, but also enhances the framework's flexibility. For instance, both the multimodal understanding and generation components can independently select their most suitable encoding methods. Experiments show that Janus surpasses previous unified model and matches or exceeds the performance of task-specific models. The simplicity, high flexibility, and effectiveness of Janus make it a strong candidate for next-generation unified multimodal models.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

deepseek-ai/janus officialmentioned in papermentioned on GitHubpytorchMIT report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Visual Question Answering

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
WISE Janus Biology 0.28 #11 of 11 Archive leaderboard report
WISE Janus Chemistry 0.14 #11 of 11 Archive leaderboard report
WISE Janus Cultural 0.16 #11 of 11 Archive leaderboard report
WISE Janus Overall 0.23 #11 of 11 Archive leaderboard report
WISE Janus Physics 0.30 #11 of 11 Archive leaderboard report
WISE Janus Space 0.35 #11 of 11 Archive leaderboard report
WISE Janus Time 0.26 #11 of 11 Archive leaderboard report
Visual Question Answering MM-Vet Janus GPT-4 score 34.3 #168 of 231 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections