{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/lingoly-too-disentangling-memorisation-from","title":"LINGOLY-TOO: Disentangling Memorisation from Reasoning with Linguistic Templatisation and Orthographic Obfuscation","arxiv_id":"2503.02972","date":"2025-03-04","proceeding":null,"authors":["Jude Khouja","Karolina Korgul","Simi Hellsten","Lingyi Yang","Vlad Neacs","Harry Mayne","Ryan Kearns","Andrew Bean","Adam Mahdi"],"abstract":"Effective evaluation of the reasoning capabilities of large language models (LLMs) are susceptible to overestimation due to data exposure of evaluation benchmarks. We introduce a framework for producing linguistic reasoning problems that reduces the effect of memorisation in model performance estimates and apply this framework to develop LINGOLY-TOO, a challenging evaluation benchmark for linguistic reasoning. By developing orthographic templates, we dynamically obfuscate the writing systems of real languages to generate numerous question variations. These variations preserve the reasoning steps required for each solution while reducing the likelihood of specific problem instances appearing in model training data. Our experiments demonstrate that frontier models, including OpenAI o1-preview and DeepSeem R1, struggle with advanced reasoning. Our analysis also shows that LLMs exhibit noticeable variance in accuracy across permutations of the same problem, and on average perform better on questions appearing in their original orthography. Our findings highlight the opaque nature of response generation in LLMs and provide evidence that prior data exposure contributes to overestimating the reasoning capabilities of frontier models.","url_abs":"https://arxiv.org/abs/2503.02972v1","url_pdf":"https://arxiv.org/pdf/2503.02972v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"response-generation","task_name":"Response Generation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2503.02972","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2503.02972"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/jkhouja/LingOly-TOO","reach":null}],"summary":{"ran_draft_wrong":5},"by_repo_kind":{"found_in_text":{"samples":5,"ran":5,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":5,"samples":[{"code_sha256_prefix":"d2527864114418bb","entry":"clean_key","repo":"jkhouja/LingOly-TOO","repo_kind":"found_in_text","path":"testing/code/scoring.py","file_url":"https://github.com/jkhouja/LingOly-TOO/blob/HEAD/testing/code/scoring.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"d2527864114418bb"}},{"code_sha256_prefix":"b815d1398561265a","entry":"extract_json_substrings","repo":"jkhouja/LingOly-TOO","repo_kind":"found_in_text","path":"testing/code/scoring.py","file_url":"https://github.com/jkhouja/LingOly-TOO/blob/HEAD/testing/code/scoring.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"b815d1398561265a"}},{"code_sha256_prefix":"e56e8b54b48d6e6f","entry":"find_match","repo":"jkhouja/LingOly-TOO","repo_kind":"found_in_text","path":"testing/code/scoring.py","file_url":"https://github.com/jkhouja/LingOly-TOO/blob/HEAD/testing/code/scoring.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"e56e8b54b48d6e6f"}},{"code_sha256_prefix":"d1ef341bc9f11b9e","entry":"load_cache","repo":"jkhouja/LingOly-TOO","repo_kind":"found_in_text","path":"testing/code/benchmark_model.py","file_url":"https://github.com/jkhouja/LingOly-TOO/blob/HEAD/testing/code/benchmark_model.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"d1ef341bc9f11b9e"}},{"code_sha256_prefix":"afdb295fa3ac671f","entry":"manipulate_prompts","repo":"jkhouja/LingOly-TOO","repo_kind":"found_in_text","path":"testing/code/benchmark_model.py","file_url":"https://github.com/jkhouja/LingOly-TOO/blob/HEAD/testing/code/benchmark_model.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"afdb295fa3ac671f"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}