Datasets › MiniF2F

MiniF2F

Introduced by Kunhao Zheng et al. in MiniF2F: a cross-system benchmark for formal Olympiad-level mathematics31 Aug 2021 archive 2025-07-28

MiniF2F is a dataset of formal Olympiad-level mathematics problems statements intended to provide a unified cross-system benchmark for neural theorem proving. The miniF2F benchmark currently targets Metamath, Lean, and Isabelle and consists of 488 problem statements drawn from the AIME, AMC, and the International Mathematical Olympiad (IMO), as well as material from high-school and undergraduate mathematics courses.

Benchmarks archive 2025-07-28

All 2 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Automated Theorem Proving miniF2F-test Kimina-Prover-Preview cumulative 80.74 Kimina-Prover Preview: Towards Large Formal Reasoning... moonshotai/kimina-prover-preview +1 29 Compare
Automated Theorem Proving miniF2F-valid Lean GPT-f Pass@8 29.3 MiniF2F: a cross-system benchmark for formal... openai/minif2f +3 10 Compare

Papers archive 2025-07-28

17 shown of 17 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 84. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Kimina-Prover Preview: Towards Large Formal Reasoning Models with Reinforcement Learning 2 1 15 Apr 2025 ran 3 of 3 samples (0 unverified; 3 pointer-only for licence)
Efficient Neural Theorem Proving via Fine-grained Proof Structure Analysis 1 1 30 Jan 2025 not harvested
SubgoalXL: Subgoal-based Expert Learning for Theorem Proving 1 1 20 Aug 2024 not harvested
DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Search 2 1 15 Aug 2024 ran 4 of 10 samples (6 unverified)
DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data 0 1 23 May 2024 not harvested
An Empirical Study of Data Ability Boundary in LLMs' Math Reasoning 1 1 23 Feb 2024 ran 9 of 11 samples (2 unverified; 11 pointer-only for licence)
Llemma: An Open Language Model For Mathematics 4 2 16 Oct 2023 ran 6 of 8 samples (2 unverified)
An In-Context Learning Agent for Formal Theorem-Proving 1 3 6 Oct 2023 not harvested
LEGO-Prover: Neural Theorem Proving with Growing Libraries 1 2 1 Oct 2023 not harvested
Lyra: Orchestrating Dual Correction in Automated Theorem Proving 1 2 27 Sep 2023 ran 7 of 8 samples (1 unverified)
Decomposing the Enigma: Subgoal-based Demonstration Learning for Formal Theorem Proving 1 1 25 May 2023 not harvested
Draft, Sketch, and Prove: Guiding Formal Theorem Provers with Informal Proofs 3 3 21 Oct 2022 not harvested
HyperTree Proof Search for Neural Theorem Proving 0 8 23 May 2022 not harvested
Thor: Wielding Hammers to Integrate Language Models and Automated Theorem Provers 0 2 22 May 2022 not harvested
Formal Mathematics Statement Curriculum Learning 1 1 3 Feb 2022 not harvested
MiniF2F: a cross-system benchmark for formal Olympiad-level mathematics 4 6 31 Aug 2021 ran 3 of 3 samples (0 unverified)
Proof Artifact Co-training for Theorem Proving with Language Models 4 1 11 Feb 2021 ran 3 of 4 samples (1 unverified)

Dataset loaders archive 2025-07-28

1 loader as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • MiniF2F
  • miniF2F-valid
  • miniF2F-test

3 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections