{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/arxiv-2606-06526","title":"CrowdMath: A Dataset of Crowdsourced Mathematical Research Discussions","arxiv_id":"2606.06526","date":"2026-06-02","proceeding":null,"authors":["Sherin Muckatira","Jesse Geneson","Slava Gerovitch","Pavel Etingof","Mikhail Gronas","Anna Rumshisky"],"abstract":"Large language models have made substantial progress on mathematical reasoning, but existing benchmarks typically evaluate well-specified problems with final answers, step-by-step solutions, or complete proofs. They do not capture collaborative open-problem solving: a setting in which participants propose partial arguments, identify gaps or errors in prior steps, repair flawed reasoning, and gradually synthesize incremental contributions into a proof. We introduce CrowdMath, a dataset of 164 expert-annotated progress chains from the MIT PRIMES--Art of Problem Solving (AoPS) CrowdMath program (2016-2025), a collaborative research initiative whose discussions have led to peer-reviewed publications. Each chain traces a multi-participant forum discussion from an open-problem statement to a completed proof. Posts are labeled by their functional roles in the evolving solution process, including partial progress, proof completion, erroneous reasoning, and error identification. We define evaluation tasks and benchmark six frontier models. Models achieve 83-88% accuracy on next-post prediction, suggesting that they can follow the local flow of mathematical discussion. However, they struggle to identify the functional significance of individual contributions with the best model achieving only 0.42 macro-F1 on post-role classification. CrowdMath exposes a gap between solving well-specified mathematical problems and understanding collaborative mathematical progress as it unfolds.","url_abs":"https://arxiv.org/abs/2606.06526","url_pdf":"https://arxiv.org/pdf/2606.06526","source":{"archive":null,"snapshot":"2025-07-28","note":"not in the Papers with Code archive (frozen at the snapshot)","row_kind":"graph","title_abstract_authors_date":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)"},"code_links":[],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2606.06526"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"mentioned_in_github":null,"is_official":null,"provenance":"deterministic:regex_extraction","mentioned_in_paper":null,"url":"https://github.com/text-machine-lab/crowdmath","reach":null}],"summary":{"ran":1,"ran_draft_wrong":4},"by_repo_kind":{"found_in_text":{"samples":5,"ran":5,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"0a4a6510fb7dcb1a","entry":"Job","repo":"text-machine-lab/crowdmath","repo_kind":"found_in_text","path":"llm_eval/harness/tasks.py","file_url":"https://github.com/text-machine-lab/crowdmath/blob/HEAD/llm_eval/harness/tasks.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"0a4a6510fb7dcb1a"}},{"code_sha256_prefix":"4a9b43eac04f845f","entry":"_build_task1_prior_context","repo":"text-machine-lab/crowdmath","repo_kind":"found_in_text","path":"llm_eval/harness/tasks.py","file_url":"https://github.com/text-machine-lab/crowdmath/blob/HEAD/llm_eval/harness/tasks.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"4a9b43eac04f845f"}},{"code_sha256_prefix":"e0c6cfd1c1696cd4","entry":"fill","repo":"text-machine-lab/crowdmath","repo_kind":"found_in_text","path":"llm_eval/harness/tasks.py","file_url":"https://github.com/text-machine-lab/crowdmath/blob/HEAD/llm_eval/harness/tasks.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"e0c6cfd1c1696cd4"}},{"code_sha256_prefix":"44d3cbde24b5ce80","entry":"load","repo":"text-machine-lab/crowdmath","repo_kind":"found_in_text","path":"llm_eval/harness/tasks.py","file_url":"https://github.com/text-machine-lab/crowdmath/blob/HEAD/llm_eval/harness/tasks.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"44d3cbde24b5ce80"}},{"code_sha256_prefix":"737db9d6379305f3","entry":"task1_iter_jobs","repo":"text-machine-lab/crowdmath","repo_kind":"found_in_text","path":"llm_eval/harness/tasks.py","file_url":"https://github.com/text-machine-lab/crowdmath/blob/HEAD/llm_eval/harness/tasks.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"737db9d6379305f3"}}]},"arxiv_metadata":{"licence":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)","fields":["title","abstract","authors","date"],"primary_category":"cs.LG","source":"arxiv_2026.jsonl"},"syntology_extracted_results":null}