{"url":"/dataset/vibe-eval","name":"Vibe-Eval","full_name":null,"description_markdown":"Vibe-Eval is a new open benchmark and framework for evaluating multimodal chat models¹². It was introduced by Reka Technologies⁴ and is designed to rigorously test these models' visual understanding capabilities⁴. Here are some key points about Vibe-Eval:\r\n\r\n- It consists of **269 ultra high-quality image-text prompts** and their ground truth responses¹.\r\n- The prompts and responses have been extensively checked multiple times by the Reka team¹.\r\n- Vibe-Eval is designed to be difficult, challenging even to the current frontier models, and to induce greater separability among frontier-class models¹.\r\n- On 50% of the hard set, all frontier models fail to arrive at a perfect answer, leaving a lot of headroom for progress¹.\r\n- The prompts are created by actual AI experts who have a strong familiarity with the performance of frontier models¹.\r\n- While MMMU has been a pretty solid standard for evaluating multimodal models, it is still fundamentally a multiple-choice benchmark¹. Vibe-Eval, on the other hand, is an open-ended evaluation setup¹.\r\n- They also discuss challenges and trade-offs between human and model-based automatic evaluation and propose a lightweight automatic evaluation protocol based on Reka Core¹.\r\n- They plan to periodically run formal human evaluations on public models that do well on this benchmark¹.\r\n\r\n(1) Vibe-Eval: A new open and hard evaluation suite for measuring progress .... https://www.reka.ai/news/vibe-eval.\r\n(2) Vibe-Eval: A hard evaluation suite for measuring progress of multimodal .... https://arxiv.org/pdf/2405.02287.\r\n(3) This AI Paper by Reka AI Introduces Vibe-Eval: A Comprehensive Suite .... https://www.marktechpost.com/2024/05/02/this-ai-paper-by-reka-ai-introduces-vibe-eval-a-comprehensive-suite-for-evaluating-ai-multimodal-models/.\r\n(4) Vibe-Eval: A hard evaluation suite for measuring progress of multimodal .... https://arxivtools.blob.core.windows.net/xueshuxiangzipaperhtml/2024_5_6/2405.02287.pdf.","description_withheld":null,"homepage":"https://github.com/reka-ai/reka-vibe-eval","introduced_date":"2024-05-03","introduced_date_note":null,"introduced_by":{"paper":"/paper/vibe-eval-a-hard-evaluation-suite-for","title":"Vibe-Eval: A hard evaluation suite for measuring progress of multimodal language models","first_author":"Piotr Padlewski","url":null},"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["Vibe-Eval"],"data_loaders":[],"num_papers_in_archive":8,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}