{"url":"/dataset/charxiv","name":"CharXiv","full_name":null,"description_markdown":"CharXiv is a comprehensive evaluation suite for testing the chart understanding capabilities of Multimodal Large Language Models (MLLMs)¹². It was proposed to address the limitations of existing datasets that often focus on oversimplified and homogeneous charts with template-based questions¹².\r\n\r\nHere are some key features of CharXiv:\r\n- It includes **2,323 natural, challenging, and diverse charts** from arXiv papers¹².\r\n- CharXiv includes two types of questions¹²:\r\n    1. **Descriptive questions** about examining basic chart elements.\r\n    2. **Reasoning questions** that require synthesizing information across complex visual elements in the chart.\r\n- All charts and questions are **handpicked, curated, and verified by human experts**¹².\r\n\r\nThe results from CharXiv reveal a substantial gap between the reasoning skills of the strongest proprietary model (i.e., GPT-4o), which achieves 47.1% accuracy, and the strongest open-source model (i.e., InternVL Chat V1.5), which achieves 29.2%². All models lag far behind human performance of 80.5%, underscoring weaknesses in the chart understanding capabilities of existing MLLMs².\r\n\r\n(1) [2406.18521] CharXiv: Charting Gaps in Realistic Chart Understanding in .... https://arxiv.org/abs/2406.18521.\r\n(2) CharXiv. https://charxiv.github.io/.\r\n(3) ChinaXiv.org 中国科学院科技论文预发布平台. https://chinaxiv.org/home.htm.\r\n(4) undefined. https://doi.org/10.48550/arXiv.2406.18521.","description_withheld":null,"homepage":"https://github.com/princeton-nlp/CharXiv","introduced_date":"2024-06-26","introduced_date_note":null,"introduced_by":{"paper":"/paper/charxiv-charting-gaps-in-realistic-chart","title":"CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs","first_author":"ZiRui Wang","url":null},"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["CharXiv"],"data_loaders":[],"num_papers_in_archive":23,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}