{"url":"/dataset/blink","name":"BlINK","full_name":null,"description_markdown":"BLINK is a new benchmark for multimodal language models (LLMs) that focuses on core visual perception abilities not found in other evaluations¹². \r\n\r\nMost of the BLINK tasks can be solved by humans “within a blink” (e.g., relative depth estimation, visual correspondence, forensics detection, and multi-view reasoning)¹². However, these perception-demanding tasks cast significant challenges for current multimodal LLMs because they resist mediation through natural language¹².\r\n\r\nBLINK reformats 14 classic computer vision tasks into 3,807 multiple-choice questions, paired with single or multiple images and visual prompting¹². While humans get 95.70% accuracy on average, BLINK is surprisingly challenging for existing multimodal LLMs: even the best-performing GPT-4V and Gemini achieve accuracies of 51.26% and 45.72%, only 13.17% and 7.63% higher than random guessing, indicating that such perception abilities have not “emerged” yet in recent multimodal LLMs¹².\r\n\r\nThe BLINK benchmark is designed to stimulate the community to help multimodal LLMs catch up with human-level visual perception¹². It includes diverse visual prompting, beyond recognition perception, and visual commonsense¹. The benchmark is available on GitHub¹.\r\n\r\n(1) zeyofu/BLINK_Benchmark - GitHub. https://github.com/zeyofu/BLINK_Benchmark.\r\n(2) BLINK: Multimodal Large Language Models Can See but Not Perceive. https://arxiv.org/abs/2404.12390.\r\n(3) Releases · zeyofu/BLINK_Benchmark · GitHub. https://github.com/zeyofu/BLINK_Benchmark/releases.\r\n(4) undefined. https://doi.org/10.48550/arXiv.2404.12390.","description_withheld":null,"homepage":"https://zeyofu.github.io/blink","introduced_date":"2024-04-18","introduced_date_note":null,"introduced_by":{"paper":"/paper/blink-multimodal-large-language-models-can","title":"BLINK: Multimodal Large Language Models Can See but Not Perceive","first_author":"Xingyu Fu","url":null},"license":null,"modalities":[],"tasks":[{"name":"Spatial Reasoning","url":"/task/spatial-reasoning","datasets_with_task":"/datasets/task/spatial-reasoning"}],"languages":[],"variants":["BlINK"],"data_loaders":[],"num_papers_in_archive":58,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/spatial-reasoning-on-blink","task":"Spatial Reasoning","dataset_variant":"BlINK","rows":0,"metrics":["Overall Success Rate"],"first_row_in_archive_order":null,"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}