{"url":"/dataset/mathverse","name":"MathVerse","full_name":null,"description_markdown":"MathVerse is an innovative benchmark specifically designed to rigorously evaluate the capabilities of Multi-modal Large Language Models (MLLMs) in interpreting and reasoning with visual information in mathematical problems. Developed by a research team from CUHK MMLab and Shanghai Artificial Intelligence Laboratory, MathVerse offers an equitable and comprehensive assessment of MLLMs' ability to understand and process visual diagrams for mathematical reasoning.\r\n\r\nThe benchmark consists of 2,612 high-quality, multi-subject math problems with diagrams, meticulously collected from publicly available sources. Each problem is then transformed by human annotators into six distinct versions, each offering varying degrees of information content in multi-modality. This approach results in a total of 15K test samples, allowing MathVerse to thoroughly examine whether and how much MLLMs can genuinely comprehend visual diagrams for solving mathematical problems.","description_withheld":null,"homepage":"https://mathverse-cuhk.github.io","introduced_date":"2024-03-21","introduced_date_note":null,"introduced_by":{"paper":"/paper/mathverse-does-your-multi-modal-llm-truly-see","title":"MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?","first_author":"Renrui Zhang","url":null},"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["MathVerse"],"data_loaders":[],"num_papers_in_archive":82,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}