{"url":"/dataset/rebus","name":"REBUS","full_name":"A Robust Evaluation Benchmark of Understanding Symbols","description_markdown":"Recent advances in large language models have led to the development of multimodal LLMs\r\n(MLLMs), which take both image data and text as an input. Virtually all of these models\r\nhave been announced within the past year, leading to a significant need for benchmarks\r\nevaluating the abilities of these models to reason truthfully and accurately on a diverse set\r\nof tasks. When Google announced Gemini (Gemini Team et al., 2023), they showcased its\r\nability to solve rebuses—wordplay puzzles which involve creatively adding and subtracting\r\nletters from words derived from text and images. The diversity of rebuses allows for a\r\nbroad evaluation of multimodal reasoning capabilities, including image recognition, multi-\r\nstep reasoning, and understanding the human creator’s intent.\r\nWe present REBUS: a collection of 333 hand-crafted rebuses spanning 13 diverse cate-\r\ngories, including hand-drawn and digital images created by nine contributors. Samples are\r\npresented in Table 1. Notably, GPT-4V, the most powerful model we evaluated, answered\r\nonly 24% of puzzles correctly, highlighting the poor capabilities of MLLMs in new and unex-\r\npected domains to which human reasoning generalizes with comparative ease. Open-source\r\nmodels perform even worse, with a median accuracy below 1%. We notice that models\r\noften give faithless explanations, fail to change their minds after an initial approach doesn’t\r\nwork, and remain highly uncalibrated on their own abilities.","description_withheld":null,"homepage":"https://cavendishlabs.org/rebus/","introduced_date":"2024-01-09","introduced_date_note":null,"introduced_by":null,"license":null,"modalities":[{"name":"Images","url":"/datasets/modality/images"},{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Logical Reasoning","url":"/task/logical-reasoning","datasets_with_task":"/datasets/task/logical-reasoning"},{"name":"Multimodal Deep Learning","url":"/task/multimodal-deep-learning","datasets_with_task":"/datasets/task/multimodal-deep-learning"},{"name":"Multimodal Reasoning","url":"/task/multimodal-reasoning","datasets_with_task":"/datasets/task/multimodal-reasoning"},{"name":"multimodal generation","url":"/task/multimodal-generation","datasets_with_task":"/datasets/task/multimodal-generation"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["REBUS"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/multimodal-reasoning-on-rebus","task":"Multimodal Reasoning","dataset_variant":"REBUS","rows":8,"metrics":["Accuracy"],"first_row_in_archive_order":{"model":"GPT-4V","paper":"/paper/rebus-a-robust-evaluation-benchmark-of-1","metrics":{"Accuracy":"24.0"},"code_links":[{"title":"cvndsh/rebus","url":"https://github.com/cvndsh/rebus"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/rebus-a-robust-evaluation-benchmark-of-1","title":"REBUS: A Robust Evaluation Benchmark of Understanding Symbols","date":"2024-01-11","rows_on_this_dataset":8,"code_links":1,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}