Browse State-of-the-Art › Object Hallucination
Object Hallucination
42 papers with code · 0 benchmarks · 0 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
No dataset record in the archive lists this task.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 42 papers with code (71 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
Mitigating Object Hallucinations in Large Vision-Language Models through Visual Contrastive Decoding28 Nov 2023 7 repositories listed Syntology ran 9 of 14 samples · 5 unverified · 1 pointer-only (licence)Large Vision-Language Models (LVLMs) have advanced considerably, intertwining visual recognition and language understanding to generate content that is not only coherent but also contextually attuned.
-
17 May 2023 6 repositories listed Syntology ran 6 of 12 samples · 6 unverified · 4 pointer-only (licence)Despite the promising progress on LVLMs, we find that LVLMs suffer from the hallucination problem, i.
-
27 May 2024 5 repositories listed Syntology ran 14 of 22 samples · 8 unverified · 8 pointer-only (licence)Traditional feedback learning for hallucination reduction relies on labor-intensive manual labeling or expensive proprietary models.
-
29 Jan 2024 3 repositories listed Syntology ran 0 of 2 samples · 2 unverified · 2 pointer-only (licence)In this work, we propose a simple yet effective training strategy MoE-Tuning for LVLMs.
-
12 Jun 2024 2 repositories listedLarge audio-language models (LALMs) enhance traditional large language models by integrating audio perception capabilities, allowing them to tackle audio-related tasks.
-
1 Mar 2024 2 repositories listed Syntology ran 7 of 12 samples · 5 unverifiedWhile large vision-language models (LVLMs) have demonstrated impressive capabilities in interpreting multi-modal contexts, they invariably suffer from object hallucinations (OH).
-
23 Feb 2024 2 repositories listed Syntology ran 11 of 17 samples · 6 unverified · 17 pointer-only (licence)Large Vision-Language Models (LVLMs) are susceptible to object hallucinations, an issue in which their generated text contains non-existent objects, greatly limiting their reliability and practicality.
-
11 Oct 2023 2 repositories listed Syntology ran 7 of 8 samples · 1 unverified · 8 pointer-only (licence)We introduce Ferret, a new Multimodal Large Language Model (MLLM) capable of understanding spatial referring of any shape or granularity within an image and accurately grounding open-vocabulary descriptions.
-
3 Oct 2023 2 repositories listed Syntology ran 5 of 9 samples · 4 unverified · 9 pointer-only (licence)Current Large Multimodal Models (LMMs) achieve remarkable progress, yet there remains significant uncertainty regarding their ability to accurately apprehend visual details, that is, in performing detailed captioning.
-
11 Jun 2025 1 repository listedOur approach leverages the semantic information embedded within vision tokens by projecting them into the text token distribution space, and dynamically selecting the most relevant vision token at each decoding step…
-
10 Jun 2025 1 repository listed Syntology ran 2 of 4 samples · 2 unverified · 4 pointer-only (licence)Despite significant advancements in Vision-Language Models (VLMs), the performance of existing VLMs remains hindered by object hallucination, a critical challenge to achieving accurate visual understanding.
-
26 May 2025 1 repository listedDespite significant advancements in Large Vision-Language Models, Object Hallucination (OH) remains a persistent challenge.
-
25 Mar 2025 1 repository listed Syntology ran 4 of 5 samples · 1 unverified · 4 pointer-only (licence)The rapid advancement of large vision-language models (LVLMs) has driven significant progress in multimodal tasks, enabling models to interpret, reason, and generate outputs across both visual and textual domains.
-
17 Mar 2025 1 repository listed Syntology ran 0 of 1 samples · 1 unverified · 1 pointer-only (licence)However, these methods present two main limitations: (1) bluntly suppressing language priors can compromise coherence and accuracy of generated content, and (2) processing contrastive inputs adds computational load,…
-
13 Mar 2025 1 repository listedIn this paper, we first conduct an in-depth exploration of LVLM internal states in relation to OH issues and discover that (1) LVLM internal states are high-specificity per-token indicators of hallucination behaviors.
-
28 Feb 2025 1 repository listed Syntology ran 3 of 11 samples · 8 unverifiedFurthermore, we propose an entropy-based noise-controlling strategy to enable the injected noise to be adaptively constrained regarding the smoothness of the similarity distribution.
-
27 Feb 2025 1 repository listedRecently, Large Vision-Language Models (LVLMs) show remarkable performance across various domains.
-
29 Dec 2024 1 repository listedLarge Vision-Language Models (LVLMs) have demonstrated remarkable performance in performing complex multimodal tasks.
-
24 Dec 2024 1 repository listedHowever, current approaches primarily rely on large foundation models in a zero-shot manner or fine-tuned models with human annotations, which limits scalability due to significant computational costs.
-
30 Oct 2024 1 repository listedThe core idea of our framework is to conduct hallucination evaluation on (object, relation, object) triplets extracted from LVLMs' responses, and thus, could be easily generalized to different vision-language tasks.
-
21 Oct 2024 1 repository listed Syntology ran 0 of 1 samples · 1 unverifiedDue to the long-term decay in RoPE, LVLMs tend to hallucinate more when relevant visual cues are distant from instruction tokens in the multimodal input sequence.
-
17 Oct 2024 1 repository listedLastly, we observe that although existing methods struggle to balance the reduction of object hallucinations with maintaining text quality, SGD demonstrates robustness in handling this challenge.
-
4 Oct 2024 1 repository listed Syntology ran 2 of 3 samples · 1 unverified · 3 pointer-only (licence)Large Vision-Language Models (LVLMs) have achieved impressive performance, yet research has pointed out a serious issue with object hallucinations within these models.
-
2 Sep 2024 1 repository listedTo analyze image representations while completely avoiding the influence of all other factors other than the image representation itself, we propose a parametric-free representation alignment metric (Pfram) that can…
-
8 Jul 2024 1 repository listed Syntology ran 7 of 8 samples · 1 unverified · 8 pointer-only (licence)Large vision language models (LVLMs) often suffer from object hallucination, producing objects not present in the given images.
-
Think Before You Act: A Two-Stage Framework for Mitigating Gender Bias Towards Vision-Language Tasks27 May 2024 1 repository listedDuring answer inference, GAMA integrates the image, generated narrative, and a task-specific question prompt to infer answers for different vision-language tasks.
-
18 Feb 2024 1 repository listed Syntology ran 0 of 3 samples · 3 unverified · 3 pointer-only (licence)In this work, we adopt the intuition that the LVLM tends to respond logically consistently for existent objects but inconsistently for hallucinated objects.
-
15 Feb 2024 1 repository listed Syntology ran 6 of 7 samples · 1 unverified · 7 pointer-only (licence)Multimodal large language models (MLLMs) have attracted increasing attention in the past few years, but they may still generate descriptions that include objects not present in the corresponding images, a phenomenon…
-
1 Feb 2024 1 repository listedWe introduce Instruction Document Visual Question Answering (iDocVQA) dataset and Large Language Document (LLaDoc) model, for training Language-Vision (LV) models for document analysis and predictions on document…
-
8 Dec 2023 1 repository listed Syntology ran 13 of 17 samples · 4 unverified · 17 pointer-only (licence)Results show that ASUKA mitigates object hallucination and improves color consistency over standard diffusion and rectified flow models and other inpainting methods.
Syntology lines on 18 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections