Papers › LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

18 Mar 2024arXiv:2403.11703archive 2025-07-28

Ruyi Xu, Yuan YAO, Zonghao Guo, Junbo Cui, Zanlin Ni, Chunjiang Ge, Tat-Seng Chua, Zhiyuan Liu, Maosong Sun, Gao Huang

Visual encoding constitutes the basis of large multimodal models (LMMs) in understanding the visual world. Conventional LMMs process images in fixed sizes and limited resolutions, while recent explorations in this direction are limited in adaptivity, efficiency, and even correctness. In this work, we first take GPT-4V and LLaVA-1.5 as representative examples and expose systematic flaws rooted in their visual encoding strategy. To address the challenges, we present LLaVA-UHD, a large multimodal model that can efficiently perceive images in any aspect ratio and high resolution. LLaVA-UHD includes three key components: (1) An image modularization strategy that divides native-resolution images into smaller variable-sized slices for efficient and extensible encoding, (2) a compression module that further condenses image tokens from visual encoders, and (3) a spatial schema to organize slice tokens for LLMs. Comprehensive experiments show that LLaVA-UHD outperforms established LMMs trained with 2-3 orders of magnitude more data on 9 benchmarks. Notably, our model built on LLaVA-1.5 336x336 supports 6 times larger (i.e., 672x1088) resolution images using only 94% inference computation, and achieves 6.4 accuracy improvement on TextVQA. Moreover, the model can be efficiently trained in academic settings, within 23 hours on 8 A100 GPUs (vs. 26 hours of LLaVA-1.5). We make the data and code publicly available at https://github.com/thunlp/LLaVA-UHD.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

thunlp/llava-uhd officialmentioned in papermentioned on GitHubpytorchApache-2.0 report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Long-Context Understanding

1 archive task tag without a task page not shown.

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Long-Context Understanding MMNeedle LLaVA-Llama-3 1 Image, 2*2 Stitching, Exact Accuracy 43.8 #5 of 12 Archive leaderboard report
Long-Context Understanding MMNeedle LLaVA-Llama-3 1 Image, 4*4 Stitching, Exact Accuracy 17.5 #5 of 12 Archive leaderboard report
Long-Context Understanding MMNeedle LLaVA-Llama-3 1 Image, 8*8 Stitching, Exact Accuracy 3.3 #5 of 12 Archive leaderboard report
Long-Context Understanding MMNeedle LLaVA-Llama-3 10 Images, 1*1 Stitching, Exact Accuracy 0 #5 of 12 Archive leaderboard report
Long-Context Understanding MMNeedle LLaVA-Llama-3 10 Images, 2*2 Stitching, Exact Accuracy 0 #5 of 12 Archive leaderboard report
Long-Context Understanding MMNeedle LLaVA-Llama-3 10 Images, 4*4 Stitching, Exact Accuracy 0 #5 of 12 Archive leaderboard report
Long-Context Understanding MMNeedle LLaVA-Llama-3 10 Images, 8*8 Stitching, Exact Accuracy 0 #5 of 12 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

1D CNN

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections