{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/lmr-bench-evaluating-llm-agent-s-ability-on","title":"LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research","arxiv_id":"2506.17335","date":"2025-06-19","proceeding":null,"authors":["Shuo Yan","Ruochen Li","Ziming Luo","Zimu Wang","Daoyang Li","Liqiang Jing","Kaiyu He","Peilin Wu","George Michalopoulos","Yue Zhang","Ziyang Zhang","Mian Zhang","Zhiyu Chen","Xinya Du"],"abstract":"Large language model (LLM) agents have demonstrated remarkable potential in advancing scientific discovery. However, their capability in the fundamental yet crucial task of reproducing code from research papers, especially in the NLP domain, remains underexplored. This task includes unique complex reasoning challenges in the intellectual synthesis of abstract concepts and the comprehension of code repositories with interdependent files. Motivated by this gap, we present LMR-BENCH, a benchmark designed to systematically evaluate the capability of LLM agents on code reproduction from Language Modeling Research. It consists of 28 code reproduction tasks derived from 23 research papers published in top-tier NLP venues over the past five years, spanning nine fundamental categories. Models are provided with a research paper, a code repository containing one or more masked functions, and instructions for implementing these functions. We conduct extensive experiments in standard prompting and LLM agent settings with state-of-the-art LLMs, evaluating the accuracy of unit tests and performing LLM-based evaluation of code correctness. Experimental results reveal that even the most advanced models still exhibit persistent limitations in scientific reasoning and code synthesis, highlighting critical gaps in LLM agents' ability to autonomously reproduce scientific research","url_abs":"https://arxiv.org/abs/2506.17335v1","url_pdf":"https://arxiv.org/pdf/2506.17335v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"lmr-bench-evaluating-llm-agent-s-ability-on","repo_url":"https://github.com/du-nlp-lab/lmr-bench","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"},{"task_slug":"large-language-model","task_name":"Large Language Model"},{"task_slug":"scientific-discovery","task_name":"scientific discovery"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2506.17335","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2506.17335"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/du-nlp-lab/lmr-bench","reach":{"status":"ok"}},{"provenance":"deterministic:regex_extraction","url":"https://github.com/du-nlp-lab/LMR-Bench","reach":{"status":"ok"}}],"summary":{"unverified":6},"by_repo_kind":{"official":{"samples":6,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":6,"samples":[{"code_sha256_prefix":"9e7ebf91e7a1facb","entry":"analyze_log","repo":"du-nlp-lab/LMR-Bench","repo_kind":"official","path":"evaluation/get_statistics.py","file_url":"https://github.com/du-nlp-lab/LMR-Bench/blob/HEAD/evaluation/get_statistics.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"9e7ebf91e7a1facb"}},{"code_sha256_prefix":"a88ce98f7db66404","entry":"file_to_string","repo":"du-nlp-lab/LMR-Bench","repo_kind":"official","path":"utils/data_process/data_utils.py","file_url":"https://github.com/du-nlp-lab/LMR-Bench/blob/HEAD/utils/data_process/data_utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"a88ce98f7db66404"}},{"code_sha256_prefix":"06dfd35e58c1b307","entry":"format_task_dict","repo":"du-nlp-lab/LMR-Bench","repo_kind":"official","path":"generation/OpenHands/evaluation/benchmarks/lmr_bench/run_infer.py","file_url":"https://github.com/du-nlp-lab/LMR-Bench/blob/HEAD/generation/OpenHands/evaluation/benchmarks/lmr_bench/run_infer.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"06dfd35e58c1b307"}},{"code_sha256_prefix":"fe28459fc508fbe0","entry":"image_exists_on_dockerhub","repo":"du-nlp-lab/LMR-Bench","repo_kind":"official","path":"utils/others/docker_push.py","file_url":"https://github.com/du-nlp-lab/LMR-Bench/blob/HEAD/utils/others/docker_push.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"fe28459fc508fbe0"}},{"code_sha256_prefix":"1fd9fab4c9315ac1","entry":"strip_code_fences","repo":"du-nlp-lab/LMR-Bench","repo_kind":"official","path":"evaluation/llm_as_a_judge_evaluation.py","file_url":"https://github.com/du-nlp-lab/LMR-Bench/blob/HEAD/evaluation/llm_as_a_judge_evaluation.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"1fd9fab4c9315ac1"}},{"code_sha256_prefix":"403d3a4f178a7569","entry":"try_pull","repo":"du-nlp-lab/LMR-Bench","repo_kind":"official","path":"evaluation/unit_test_evaluation.py","file_url":"https://github.com/du-nlp-lab/LMR-Bench/blob/HEAD/evaluation/unit_test_evaluation.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"403d3a4f178a7569"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}