{"url":"/dataset/object-halbench","name":"Object HalBench","full_name":null,"description_markdown":"Object HalBench is a benchmark used to evaluate the performance of Language Models, particularly those that are multimodal (i.e., they can process and generate both text and images). It's designed to test how well these models can avoid \"hallucinations\" - generating text that is not factually grounded in the images they're processing¹.\r\n\r\nFor instance, the OmniLMM-12B model, which is a state-of-the-art open-source Language Model, has been reported to outperform GPT-4V on the Object HalBench¹. This model is aligned via a technique called multimodal RLHF (Reinforcement Learning from Human Feedback) for trustworthy behavior¹. This means it's designed to generate outputs that are more reliable and factually accurate, particularly when dealing with multimodal inputs¹. \r\n\r\n(1) openbmb/OmniLMM-12B · Hugging Face. https://huggingface.co/openbmb/OmniLMM-12B.\r\n(2) OmniLMM：准确、高效的开源多模态大模型 - 知乎. https://zhuanlan.zhihu.com/p/681251797.\r\n(3) RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from - arXiv.org. https://arxiv.org/html/2312.00849v2.\r\n(4) README.md · openbmb/MiniCPM-V-2 at main - Hugging Face. https://huggingface.co/openbmb/MiniCPM-V-2/blob/main/README.md.\r\n(5) undefined. https://github.com/OpenBMB/OmniLMM.git.","description_withheld":null,"homepage":"","introduced_date":"2018-09-06","introduced_date_note":null,"introduced_by":{"paper":"/paper/object-hallucination-in-image-captioning","title":"Object Hallucination in Image Captioning","first_author":"Anna Rohrbach","url":null},"license":null,"modalities":[],"tasks":[{"name":"Image Captioning","url":"/task/image-captioning","datasets_with_task":"/datasets/task/image-captioning"}],"languages":[],"variants":["Object HalBench"],"data_loaders":[],"num_papers_in_archive":11,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/image-captioning-on-object-halbench","task":"Image Captioning","dataset_variant":"Object HalBench","rows":3,"metrics":["chair_i","chair_s"],"first_row_in_archive_order":{"model":"RLHF-V","paper":"/paper/rlhf-v-towards-trustworthy-mllms-via-behavior","metrics":{"chair_i":"7.5","chair_s":"12.2"},"code_links":[{"title":"openbmb/minicpm-v","url":"https://github.com/openbmb/minicpm-v"},{"title":"rlhf-v/rlhf-v","url":"https://github.com/rlhf-v/rlhf-v"},{"title":"tidedra/vl-rlhf","url":"https://github.com/tidedra/vl-rlhf"},{"title":"exgc/r1v-free","url":"https://github.com/exgc/r1v-free"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/rlaif-v-aligning-mllms-through-open-source-ai","title":"RLAIF-V: Open-Source AI Feedback Leads to Super GPT-4V Trustworthiness","date":"2024-05-27","rows_on_this_dataset":2,"code_links":5,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":22,"samples_ran":14,"samples_unverified":8,"pointer_only_for_licence":8,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/rlhf-v-towards-trustworthy-mllms-via-behavior","title":"RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedback","date":"2023-12-01","rows_on_this_dataset":1,"code_links":4,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":1,"samples_harvested":22,"samples_ran":14,"samples_unverified":8,"pointer_only_for_licence":8,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}