{"url":"/dataset/mathequiv","name":"MathEquiv","full_name":"mathematical statement equivalence","description_markdown":"MathEquiv dataset is accompanied to [EquivPruner](https://github.com/Lolo1222/EquivPruner) . It is specifically designed for **mathematical statement equivalence** , serving as a versatile resource applicable to a variety of mathematical tasks and scenarios. It consists of almost 100k math sentences pair with equivalence result and reasoning step generated by GPT-4O.\r\n\r\nThe dataset consists of three splits:\r\n\r\n- `train` with 77.6k problems for training.\r\n- `test` with 9.83k samples for testing.\r\n- `valid` with 9.75k samples for validation.\r\n\r\nWe implemented a five-tiered classification system. This granular approach was adopted to enhance the stability of the GPT model's outputs, as preliminary experiments with binary classification (equivalent/non-equivalent) revealed inconsistencies in judgments. The five-tiered system yielded significantly more consistent and reliable assessments:\r\n\r\n- Level 4 (Exactly Equivalent): The statements are mathematically interchangeable in all respects, exhibiting identical meaning and form.\r\n- Level 3 (Likely Equivalent): Minor syntactic differences may be present, but the core mathematical content and logic align.\r\n- Level 2 (Indeterminable): Insufficient information is available to make a definitive judgment regarding equivalence.\r\n- Level 1 (Unlikely Equivalent): While some partial agreement may exist, critical discrepancies in logic, definition, or mathematical structure are observed.\r\n- Level 0 (Not Equivalent): The statements are fundamentally distinct in their mathematical meaning, derivation, or resultant outcomes.","description_withheld":null,"homepage":"https://huggingface.co/datasets/Jiawei1222/MathEquiv","introduced_date":"2025-05-22","introduced_date_note":null,"introduced_by":{"paper":"/paper/equivpruner-boosting-efficiency-and-quality-1","title":"EquivPruner: Boosting Efficiency and Quality in LLM-Based Search via Action Pruning","first_author":"Jiawei Liu","url":null},"license":{"name":"Apache 2.0","url":null},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"task","url":null,"datasets_with_task":"/datasets/task/task"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["MathEquiv"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-25T09:33:49+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}