Papers › Omni-SMoLA: Boosting Generalist Multimodal Models with Soft Mixture of Low-rank Experts

Omni-SMoLA: Boosting Generalist Multimodal Models with Soft Mixture of Low-rank Experts

1 Dec 2023CVPR 2024 1arXiv:2312.00968archive 2025-07-28

Jialin Wu, Xia Hu, Yaqing Wang, Bo Pang, Radu Soricut

Large multi-modal models (LMMs) exhibit remarkable performance across numerous tasks. However, generalist LMMs often suffer from performance degradation when tuned over a large collection of tasks. Recent research suggests that Mixture of Experts (MoE) architectures are useful for instruction tuning, but for LMMs of parameter size around O(50-100B), the prohibitive cost of replicating and storing the expert models severely limits the number of experts we can use. We propose Omni-SMoLA, an architecture that uses the Soft MoE approach to (softly) mix many multimodal low rank experts, and avoids introducing a significant number of new parameters compared to conventional MoE models. The core intuition here is that the large model provides a foundational backbone, while different lightweight experts residually learn specialized knowledge, either per-modality or multimodally. Extensive experiments demonstrate that the SMoLA approach helps improve the generalist performance across a broad range of generative vision-and-language tasks, achieving new SoTA generalist performance that often matches or outperforms single specialized LMM baselines, as well as new SoTA specialist performance.

PaperPDFConference PDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Chart Question AnsweringDocument AIImage CaptioningMixture-of-ExpertsObject CountingVisual Question Answering (VQA)

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Chart Question Answering ChartQA SMoLA-PaLI-X Specialist Model 1:1 Accuracy 74.6 #7 of 27 Archive leaderboard report
Chart Question Answering ChartQA SMoLA-PaLI-X Generalist Model 1:1 Accuracy 73.8 #8 of 27 Archive leaderboard report
Object Counting TallyQA-Complex SMoLA-PaLI-X Specialist Accuracy 77.1 #1 of 6 Archive leaderboard report
Object Counting TallyQA-Complex SMoLA-PaLI-X Generalist (0 shot) Accuracy 70.7 #3 of 6 Archive leaderboard report
Object Counting TallyQA-Simple SMoLA-PaLI-X Specialist Accuracy 86.3 #1 of 6 Archive leaderboard report
Object Counting TallyQA-Simple SMoLA-PaLI-X Generalist (0 shot) Accuracy 83.3 #3 of 6 Archive leaderboard report
Visual Question Answering (VQA) A-OKVQA SMoLA-PaLI-X Specialist Model DA VQA Score 70.55 #1 of 15 Archive leaderboard report
Visual Question Answering (VQA) A-OKVQA SMoLA-PaLI-X Specialist Model MC Accuracy 83.75 #1 of 15 Archive leaderboard report
Visual Question Answering (VQA) AI2D SMoLA-PaLI-X Specialist Model EM 82.5 #1 of 4 Archive leaderboard report
Visual Question Answering (VQA) AI2D SMoLA-PaLI-X Generalist Model EM 81.4 #2 of 4 Archive leaderboard report
Visual Question Answering (VQA) DocVQA test SMoLA-PaLI-X Specialist ANLS 0.908 #3 of 33 Archive leaderboard report
Visual Question Answering (VQA) DocVQA test SMoLA-PaLI-X Generalist ANLS 0.906 #4 of 33 Archive leaderboard report
Visual Question Answering (VQA) InfographicVQA SMoLA-PaLI-X Specialist ANLS 66.2 #2 of 21 Archive leaderboard report
Visual Question Answering (VQA) InfographicVQA SMoLA-PaLI-X Generalist ANLS 65.6 #4 of 21 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections