Browse State-of-the-Art › Image Comprehension
Image Comprehension
25 papers with code · 0 benchmarks · 1 dataset archive 2025-07-28
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
1 dataset whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
25 shown of 25 papers with code (49 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
26 Sep 2023 3 repositories listedWe propose InternLM-XComposer, a vision-language large model that enables advanced image-text comprehension and composition.
-
24 May 2024 2 repositories listed Syntology ran 5 of 5 samples · 0 unverifiedIn this paper, we propose SIMA, a framework that enhances visual and language modality alignment through self-improvement, eliminating the needs for external models or data.
-
27 Mar 2024 2 repositories listed Syntology ran 7 of 8 samples · 1 unverifiedWe try to narrow the gap by mining the potential of VLMs for better performance and any-to-any workflow from three aspects, i.
-
27 Feb 2025 1 repository listedTo address fine-grained compositional REC, we propose novel methods based on a Specialist-MLLM collaboration framework, leveraging the complementary strengths of them: Specialist Models handle simpler tasks efficiently,…
-
1 Jan 2025 1 repository listedMultimodal large language models (MLLMs) demonstrate strong capabilities in multimodal understanding, reasoning, and interaction but still face the fundamental limitation of hallucinations, where they generate erroneous…
-
7 Dec 2024 1 repository listedTo enhance the model's ability to capture visual information at different levels without increasing model size, we design a novel architecture called Granularity-oriented Mixture of Experts to constraint the model to…
-
5 Dec 2024 1 repository listedWe posit that if a video diffusion model can effectively de-noise video clips by taking the features of a video tokenizer as the condition, then the tokenizer has successfully captured robust spatial and temporal…
-
21 Nov 2024 1 repository listedFurthermore, we introduce MMGenBench-Test, a comprehensive benchmark developed to evaluate LMMs across 13 distinct image patterns, and MMGenBench-Domain, targeting the performance evaluation of LMMs within the…
-
19 Nov 2024 1 repository listedAs an essential visual attribute, image complexity affects human image comprehension and directly influences the performance of computer vision tasks.
-
13 Nov 2024 1 repository listedRecent multimodal retrieval methods have endowed text-based retrievers with multimodal capabilities by utilizing pre-training strategies for visual-text alignment.
-
6 Nov 2024 1 repository listedIn this paper, we introduce StreamingBench, the first comprehensive benchmark designed to evaluate the streaming video understanding capabilities of MLLMs.
-
16 Oct 2024 1 repository listedBased on this, we introduce the Flow Text with Image Insertion Benchmark (FTII-Bench), which includes 318 high-quality Chinese image-text news articles and 307 high-quality English image-text news articles, covering 10…
-
23 Sep 2024 1 repository listed Syntology ran 6 of 6 samples · 0 unverified · 6 pointer-only (licence)Referring Expression Comprehension (REC) is a crucial cross-modal task that objectively evaluates the capabilities of language understanding, image comprehension, and language-to-image grounding.
-
5 Aug 2024 1 repository listedTo fill this gap, we introduce the Multimodal Multi-image Understanding (MMIU) benchmark, a comprehensive evaluation suite designed to assess LVLMs across a wide range of multi-image tasks.
-
31 Jul 2024 1 repository listed Syntology ran 5 of 7 samples · 2 unverifiedHowever, the scale disparity between vision encoder and language model may led to LLMs assuming a predominant role in multi-modal comprehension.
-
3 Jul 2024 1 repository listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)This long-context capability allows IXC-2.
-
20 Jun 2024 1 repository listedWe first construct a Vision Question Answering (VQA) dataset of 63.
-
30 May 2024 1 repository listed Syntology ran 3 of 4 samples · 1 unverified · 1 pointer-only (licence)To further self-improve reasoning on the extracted visual information, we let the model reuse a small portion of existing instruction-tuning data and append its self-generated image descriptions to the prompts.
-
7 Apr 2024 1 repository listedThis highlights the challenging nature of our benchmark for existing models and the significant gap between the multimodal reasoning capabilities of current models and humans.
-
30 Jan 2024 1 repository listedMulti-modal large language models (MLLMs) have demonstrated remarkable success in vision and visual-language tasks within the natural image domain.
-
CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image Inputs5 Jan 2024 1 repository listedWhen exploring the development of Artificial General Intelligence (AGI), a critical task for these models involves interpreting and processing information from multiple image inputs.
-
3 Aug 2023 1 repository listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)To this end, we propose to extract features corresponding to regional objects as soft prompts for LLM, which provides a straightforward and scalable approach and eliminates the need for LLM fine-tuning.
-
3 Jul 2023 1 repository listedOn our dataset, we have devised four benchmarks to assess the performance of generated image comprehension in relation to both content and style interpretation.
-
3 Jul 2023 1 repository listed Syntology ran 2 of 4 samples · 2 unverifiedOpen-vocabulary image segmentation aims to partition an image into semantic regions according to arbitrary text descriptions.
-
12 May 2023 1 repository listedHowever, a grand challenge of exploiting LLMs for multimodal learning is the size of pre-trained LLMs which are always with billions of parameters.
Syntology lines on 8 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections