{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/llm-srbench-a-new-benchmark-for-scientific","title":"LLM-SRBench: A New Benchmark for Scientific Equation Discovery with Large Language Models","arxiv_id":"2504.10415","date":"2025-04-14","proceeding":null,"authors":["Parshin Shojaee","Ngoc-Hieu Nguyen","Kazem Meidani","Amir Barati Farimani","Khoa D Doan","Chandan K Reddy"],"abstract":"Scientific equation discovery is a fundamental task in the history of scientific progress, enabling the derivation of laws governing natural phenomena. Recently, Large Language Models (LLMs) have gained interest for this task due to their potential to leverage embedded scientific knowledge for hypothesis generation. However, evaluating the true discovery capabilities of these methods remains challenging, as existing benchmarks often rely on common equations that are susceptible to memorization by LLMs, leading to inflated performance metrics that do not reflect discovery. In this paper, we introduce LLM-SRBench, a comprehensive benchmark with 239 challenging problems across four scientific domains specifically designed to evaluate LLM-based scientific equation discovery methods while preventing trivial memorization. Our benchmark comprises two main categories: LSR-Transform, which transforms common physical models into less common mathematical representations to test reasoning beyond memorized forms, and LSR-Synth, which introduces synthetic, discovery-driven problems requiring data-driven reasoning. Through extensive evaluation of several state-of-the-art methods, using both open and closed LLMs, we find that the best-performing system so far achieves only 31.5% symbolic accuracy. These findings highlight the challenges of scientific equation discovery, positioning LLM-SRBench as a valuable resource for future research.","url_abs":"https://arxiv.org/abs/2504.10415v1","url_pdf":"https://arxiv.org/pdf/2504.10415v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"llm-srbench-a-new-benchmark-for-scientific","repo_url":"https://github.com/deep-symbolic-mathematics/llm-sr","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"llm-srbench-a-new-benchmark-for-scientific","repo_url":"https://github.com/deep-symbolic-mathematics/llm-srbench","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"equation-discovery","task_name":"Equation Discovery"},{"task_slug":"memorization","task_name":"Memorization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2504.10415","atlas_url":"https://app.syntology.ai/?focus=2504.10415","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2504.10415"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/PingchuanMa/SGA","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/deep-symbolic-mathematics/llm-srbench","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/deep-symbolic-mathematics/llm-sr","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran":4},"by_repo_kind":{"official":{"samples":4,"ran":4,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"e120137f6f7e0ddd","entry":"add_numba_decorator","repo":"deep-symbolic-mathematics/llm-sr","repo_kind":"official","path":"llmsr/evaluator_accelerate.py","file_url":"https://github.com/deep-symbolic-mathematics/llm-sr/blob/HEAD/llmsr/evaluator_accelerate.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"e120137f6f7e0ddd"}},{"code_sha256_prefix":"026999cb70a7da74","entry":"rename_function_calls","repo":"deep-symbolic-mathematics/llm-sr","repo_kind":"official","path":"llmsr/code_manipulation.py","file_url":"https://github.com/deep-symbolic-mathematics/llm-sr/blob/HEAD/llmsr/code_manipulation.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"026999cb70a7da74"}},{"code_sha256_prefix":"6ef6f4218d0a078c","entry":"text_to_function","repo":"deep-symbolic-mathematics/llm-sr","repo_kind":"official","path":"llmsr/code_manipulation.py","file_url":"https://github.com/deep-symbolic-mathematics/llm-sr/blob/HEAD/llmsr/code_manipulation.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"6ef6f4218d0a078c"}},{"code_sha256_prefix":"c8483749d376508b","entry":"text_to_program","repo":"deep-symbolic-mathematics/llm-sr","repo_kind":"official","path":"llmsr/code_manipulation.py","file_url":"https://github.com/deep-symbolic-mathematics/llm-sr/blob/HEAD/llmsr/code_manipulation.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"c8483749d376508b"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}