Datasets › Object HalBench

Object HalBench

Introduced by Anna Rohrbach et al. in Object Hallucination in Image Captioning6 Sep 2018 archive 2025-07-28

Object HalBench is a benchmark used to evaluate the performance of Language Models, particularly those that are multimodal (i.e., they can process and generate both text and images). It's designed to test how well these models can avoid "hallucinations" - generating text that is not factually grounded in the images they're processing¹.

For instance, the OmniLMM-12B model, which is a state-of-the-art open-source Language Model, has been reported to outperform GPT-4V on the Object HalBench¹. This model is aligned via a technique called multimodal RLHF (Reinforcement Learning from Human Feedback) for trustworthy behavior¹. This means it's designed to generate outputs that are more reliable and factually accurate, particularly when dealing with multimodal inputs¹.

(1) openbmb/OmniLMM-12B · Hugging Face. https://huggingface.co/openbmb/OmniLMM-12B. (2) OmniLMM:准确、高效的开源多模态大模型 - 知乎. https://zhuanlan.zhihu.com/p/681251797. (3) RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from - arXiv.org. https://arxiv.org/html/2312.00849v2. (4) README.md · openbmb/MiniCPM-V-2 at main - Hugging Face. https://huggingface.co/openbmb/MiniCPM-V-2/blob/main/README.md. (5) undefined. https://github.com/OpenBMB/OmniLMM.git.

Benchmarks archive 2025-07-28

All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Image Captioning Object HalBench RLHF-V chair_i 7.5 RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment... openbmb/minicpm-v +3 3 Compare

Papers archive 2025-07-28

2 shown of 2 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 11. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
RLAIF-V: Open-Source AI Feedback Leads to Super GPT-4V Trustworthiness 5 2 27 May 2024 ran 14 of 22 samples (8 unverified; 8 pointer-only for licence)
RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedback 4 1 1 Dec 2023 not harvested

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

No modality tagged.

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • Object HalBench

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections