{"url":"/dataset/visit-bench","name":"VisIT-Bench","full_name":null,"description_markdown":"VisIT-Bench is a new vision-language instruction following benchmark inspired by real-world use cases. Testing 70 diverse “wish-list” skills with an automated ranking system, it advances the ongoing assessment of multimodal chatbot performance.\r\n\r\n\r\nWhy VisIT-Bench 🤔?\r\nThough recent VLMs have shown promise in following instructions, their evaluation for real-world human-chatbot instructions is often limited. Typically, VLMs are evaluated through qualitative comparison of outputs, which makes it challenging to quantify progress and potential shortcomings. VisIT-Bench helps address this problem by offering a comprehensive testbed for measuring model performance across a diverse set of instruction-following tasks, inspired by real world scenarios.🌍","description_withheld":null,"homepage":"https://visit-bench.github.io/","introduced_date":"2023-08-12","introduced_date_note":null,"introduced_by":{"paper":"/paper/visit-bench-a-benchmark-for-vision-language","title":"VisIT-Bench: A Benchmark for Vision-Language Instruction Following Inspired by Real-World Use","first_author":"Yonatan Bitton","url":null},"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["VisIT-Bench"],"data_loaders":[{"repo":"https://github.com/huggingface/datasets","url":"https://huggingface.co/datasets/mlfoundations/VisIT-Bench","frameworks":["tf","pytorch","jax"]}],"num_papers_in_archive":14,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}