{"url":"/dataset/cii-bench","name":"CII-Bench","full_name":"Chinese Image Implication understanding Benchmark","description_markdown":"We introduce the **C**hinese **I**mage **I**mplication Understanding **Bench**mark **CII-Bench**, a new benchmark measuring the higher-order perceptual, reasoning and comprehension abilities of MLLMs when presented with complex Chinese implication images. These images, including abstract artworks, comics and posters, possess visual implications that require an understanding of visual details and reasoning ability. CII-Bench reveals whether current MLLMs, leveraging their inherent comprehension abilities, can accurately decode the metaphors embedded within the complex and abstract information presented in these images. \r\n\r\nCII-Bench comprises **698** Chinese images, each accompanied by 1 to 3 multiple-choice questions, totaling **800** questions. CII-Bench encompasses images from six distinct domains: Life, Art, Society, Environment, Politics, and Chinese Traditional Culture. It also features a diverse array of image types, including Illustrations, Memes, Posters, Multi-panel Comics, Single-panel Comics, and Paintings.","description_withheld":null,"homepage":"https://huggingface.co/datasets/m-a-p/CII-Bench","introduced_date":"2024-10-17","introduced_date_note":null,"introduced_by":{"paper":"/paper/can-mllms-understand-the-deep-implication","title":"Can MLLMs Understand the Deep Implication Behind Chinese Images?","first_author":"Chenhao Zhang","url":null},"license":{"name":"Apache license 2.0","url":"https://huggingface.co/datasets/choosealicense/licenses/blob/main/markdown/apache-2.0.md"},"modalities":[{"name":"Images","url":"/datasets/modality/images"},{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Question Answering","url":"/task/question-answering","datasets_with_task":"/datasets/task/question-answering"},{"name":"Visual Question Answering (VQA)","url":"/task/visual-question-answering","datasets_with_task":"/datasets/task/visual-question-answering"},{"name":"Visual Question Answering","url":"/task/visual-question-answering-1","datasets_with_task":"/datasets/task/visual-question-answering-1"},{"name":"Image to text","url":"/task/image-to-text","datasets_with_task":"/datasets/task/image-to-text"}],"languages":[{"name":"Chinese","url":"/datasets/language/chinese"}],"variants":["CII-Bench"],"data_loaders":[],"num_papers_in_archive":2,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}