{"url":"/dataset/mathvista","name":"MathVista","full_name":"Mathematical Reasoning of in Visual Contexts","description_markdown":"**MathVista** is a consolidated Mathematical reasoning benchmark within Visual contexts. It consists of **three newly created datasets, IQTest, FunctionQA, and PaperQA**, which address the missing visual domains and are tailored to evaluate logical reasoning on puzzle test figures, algebraic reasoning over functional plots, and scientific reasoning with academic paper figures, respectively. It also incorporates **9 MathQA datasets** and **19 VQA datasets** from the literature, which significantly enrich the diversity and complexity of visual perception and mathematical reasoning challenges within our benchmark. In total, **MathVista** includes **6,141 examples** collected from **31 different datasets**.\r\n\r\n- Project: [https://mathvista.github.io/](https://mathvista.github.io/)\r\n- Visualization: [https://mathvista.github.io/#visualization](https://mathvista.github.io/#visualization)\r\n- Leaderboard: [https://mathvista.github.io/#leaderboard](https://mathvista.github.io/#leaderboard)\r\n- Paper: [https://arxiv.org/abs/2310.02255](https://arxiv.org/abs/2310.02255)\r\n- Data: [https://huggingface.co/datasets/AI4Math/MathVista](https://huggingface.co/datasets/AI4Math/MathVista)\r\n- Code: [https://github.com/lupantech/MathVista](https://github.com/lupantech/MathVista)","description_withheld":null,"homepage":"https://mathvista.github.io/","introduced_date":"2023-10-03","introduced_date_note":null,"introduced_by":{"paper":"/paper/mathvista-evaluating-mathematical-reasoning","title":"MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts","first_author":"Pan Lu","url":null},"license":{"name":"CC BY-SA 4.0","url":"https://creativecommons.org/licenses/by-sa/4.0/"},"modalities":[{"name":"Images","url":"/datasets/modality/images"},{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Question Answering","url":"/task/question-answering","datasets_with_task":"/datasets/task/question-answering"},{"name":"Visual Question Answering (VQA)","url":"/task/visual-question-answering","datasets_with_task":"/datasets/task/visual-question-answering"},{"name":"Visual Question Answering","url":"/task/visual-question-answering-1","datasets_with_task":"/datasets/task/visual-question-answering-1"},{"name":"Mathematical Reasoning","url":"/task/mathematical-reasoning","datasets_with_task":"/datasets/task/mathematical-reasoning"},{"name":"Visual Reasoning","url":"/task/visual-reasoning","datasets_with_task":"/datasets/task/visual-reasoning"},{"name":"Math Word Problem Solving","url":"/task/math-word-problem-solving","datasets_with_task":"/datasets/task/math-word-problem-solving"},{"name":"Multiple-choice","url":"/task/multiple-choice","datasets_with_task":"/datasets/task/multiple-choice"}],"languages":[{"name":"English","url":"/datasets/language/english"},{"name":"Chinese","url":"/datasets/language/chinese"},{"name":"Persian","url":"/datasets/language/persian"}],"variants":["MathVista"],"data_loaders":[{"repo":"https://github.com/huggingface/datasets","url":"https://huggingface.co/datasets/JierunChen/MathVista_with_difficulty_level","frameworks":["tf","pytorch","jax"]},{"repo":"https://github.com/huggingface/datasets","url":"https://huggingface.co/datasets/AI4Math/MathVista","frameworks":["tf","pytorch","jax"]},{"repo":"https://github.com/lupantech/MathVista","url":"https://huggingface.co/datasets/AI4Math/MathVista","frameworks":["pytorch"]}],"num_papers_in_archive":242,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}