{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/vibe-eval-a-hard-evaluation-suite-for","title":"Vibe-Eval: A hard evaluation suite for measuring progress of multimodal language models","arxiv_id":"2405.02287","date":"2024-05-03","proceeding":null,"authors":["Piotr Padlewski","Max Bain","Matthew Henderson","Zhongkai Zhu","Nishant Relan","Hai Pham","Donovan Ong","Kaloyan Aleksiev","Aitor Ormazabal","Samuel Phua","Ethan Yeo","Eugenie Lamprecht","Qi Liu","Yuqi Wang","Eric Chen","Deyu Fu","Lei LI","Che Zheng","Cyprien de Masson d'Autume","Dani Yogatama","Mikel Artetxe","Yi Tay"],"abstract":"We introduce Vibe-Eval: a new open benchmark and framework for evaluating multimodal chat models. Vibe-Eval consists of 269 visual understanding prompts, including 100 of hard difficulty, complete with gold-standard responses authored by experts. Vibe-Eval is open-ended and challenging with dual objectives: (i) vibe checking multimodal chat models for day-to-day tasks and (ii) rigorously testing and probing the capabilities of present frontier models. Notably, our hard set contains >50% questions that all frontier models answer incorrectly. We explore the nuances of designing, evaluating, and ranking models on ultra challenging prompts. We also discuss trade-offs between human and automatic evaluation, and show that automatic model evaluation using Reka Core roughly correlates to human judgment. We offer free API access for the purpose of lightweight evaluation and plan to conduct formal human evaluations for public models that perform well on the Vibe-Eval's automatic scores. We release the evaluation code and data, see https://github.com/reka-ai/reka-vibe-eval","url_abs":"https://arxiv.org/abs/2405.02287v1","url_pdf":"https://arxiv.org/pdf/2405.02287v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"vibe-eval-a-hard-evaluation-suite-for","repo_url":"https://github.com/reka-ai/reka-vibe-eval","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[],"methods":[{"method_slug":"set","method_name":"SET"}],"datasets_introduced":[{"slug":"vibe-eval","name":"Vibe-Eval","full_name":""}],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2405.02287","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2405.02287"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/reka-ai/reka-vibe-eval","reach":null}],"summary":{"ran_draft_wrong":1,"unverified":2},"by_repo_kind":{"official":{"samples":3,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"3e254180b70e9e06","entry":"make_evaluator_prompt","repo":"reka-ai/reka-vibe-eval","repo_kind":"official","path":"evaluate.py","file_url":"https://github.com/reka-ai/reka-vibe-eval/blob/HEAD/evaluate.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"3e254180b70e9e06"}},{"code_sha256_prefix":"330f7dfe33fdca3e","entry":"evaluate","repo":"reka-ai/reka-vibe-eval","repo_kind":"official","path":"evaluate.py","file_url":"https://github.com/reka-ai/reka-vibe-eval/blob/HEAD/evaluate.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"330f7dfe33fdca3e"}},{"code_sha256_prefix":"5d5e85cc5f6e22f9","entry":"get_model","repo":"reka-ai/reka-vibe-eval","repo_kind":"official","path":"models/generate.py","file_url":"https://github.com/reka-ai/reka-vibe-eval/blob/HEAD/models/generate.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"5d5e85cc5f6e22f9"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}