{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/lingoly-a-benchmark-of-olympiad-level","title":"LINGOLY: A Benchmark of Olympiad-Level Linguistic Reasoning Puzzles in Low-Resource and Extinct Languages","arxiv_id":"2406.06196","date":"2024-06-10","proceeding":null,"authors":["Andrew M. Bean","Simi Hellsten","Harry Mayne","Jabez Magomere","Ethan A. Chi","Ryan Chi","Scott A. Hale","Hannah Rose Kirk"],"abstract":"In this paper, we present the LingOly benchmark, a novel benchmark for advanced reasoning abilities in large language models. Using challenging Linguistic Olympiad puzzles, we evaluate (i) capabilities for in-context identification and generalisation of linguistic patterns in very low-resource or extinct languages, and (ii) abilities to follow complex task instructions. The LingOly benchmark covers more than 90 mostly low-resource languages, minimising issues of data contamination, and contains 1,133 problems across 6 formats and 5 levels of human difficulty. We assess performance with both direct accuracy and comparison to a no-context baseline to penalise memorisation. Scores from 11 state-of-the-art LLMs demonstrate the benchmark to be challenging, and models perform poorly on the higher difficulty problems. On harder problems, even the top model only achieved 38.7% accuracy, a 24.7% improvement over the no-context baseline. Large closed models typically outperform open models, and in general, the higher resource the language, the better the scores. These results indicate, in absence of memorisation, true multi-step out-of-domain reasoning remains a challenge for current language models.","url_abs":"https://arxiv.org/abs/2406.06196v3","url_pdf":"https://arxiv.org/pdf/2406.06196v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"lingoly-a-benchmark-of-olympiad-level","repo_url":"https://github.com/am-bean/lingOly","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"logical-reasoning","task_name":"Logical Reasoning"}],"methods":[],"datasets_introduced":[{"slug":"lingoly","name":"LingOly","full_name":""}],"methods_introduced":[],"results":[{"leaderboard":"/sota/logical-reasoning-on-lingoly","task":"Logical Reasoning","dataset":"LingOly","model":"Claude Opus","rank_in_archive_order":1,"of":11,"metrics":{"Delta_NoContext":"28.8%","Exact Match Accuracy":"46.3%"},"uses_additional_data":false},{"leaderboard":"/sota/logical-reasoning-on-lingoly","task":"Logical Reasoning","dataset":"LingOly","model":"GPT-4o","rank_in_archive_order":2,"of":11,"metrics":{"Delta_NoContext":"25.1%","Exact Match Accuracy":"37.6%"},"uses_additional_data":false},{"leaderboard":"/sota/logical-reasoning-on-lingoly","task":"Logical Reasoning","dataset":"LingOly","model":"Gemini 1.5 Pro","rank_in_archive_order":3,"of":11,"metrics":{"Delta_NoContext":"23.4%","Exact Match Accuracy":"32.1%"},"uses_additional_data":false},{"leaderboard":"/sota/logical-reasoning-on-lingoly","task":"Logical Reasoning","dataset":"LingOly","model":"GPT-4","rank_in_archive_order":4,"of":11,"metrics":{"Delta_NoContext":"21.5%","Exact Match Accuracy":"33.4%"},"uses_additional_data":false},{"leaderboard":"/sota/logical-reasoning-on-lingoly","task":"Logical Reasoning","dataset":"LingOly","model":"Command R+","rank_in_archive_order":5,"of":11,"metrics":{"Delta_NoContext":"11.6%","Exact Match Accuracy":"21.5%"},"uses_additional_data":false},{"leaderboard":"/sota/logical-reasoning-on-lingoly","task":"Logical Reasoning","dataset":"LingOly","model":"GPT-3.5","rank_in_archive_order":6,"of":11,"metrics":{"Delta_NoContext":"11.2%","Exact Match Accuracy":"21.2%"},"uses_additional_data":false},{"leaderboard":"/sota/logical-reasoning-on-lingoly","task":"Logical Reasoning","dataset":"LingOly","model":"Mixtral 8x7B","rank_in_archive_order":7,"of":11,"metrics":{"Delta_NoContext":"6.4%","Exact Match Accuracy":"14.2%"},"uses_additional_data":false},{"leaderboard":"/sota/logical-reasoning-on-lingoly","task":"Logical Reasoning","dataset":"LingOly","model":"Llama 3 8B","rank_in_archive_order":8,"of":11,"metrics":{"Delta_NoContext":"4.9%","Exact Match Accuracy":"11.4%"},"uses_additional_data":false},{"leaderboard":"/sota/logical-reasoning-on-lingoly","task":"Logical Reasoning","dataset":"LingOly","model":"Llama 3 70B","rank_in_archive_order":9,"of":11,"metrics":{"Delta_NoContext":"2.9%","Exact Match Accuracy":"10.3%"},"uses_additional_data":false},{"leaderboard":"/sota/logical-reasoning-on-lingoly","task":"Logical Reasoning","dataset":"LingOly","model":"Gemma 7B","rank_in_archive_order":10,"of":11,"metrics":{"Delta_NoContext":"2.2%","Exact Match Accuracy":"4.9%"},"uses_additional_data":false},{"leaderboard":"/sota/logical-reasoning-on-lingoly","task":"Logical Reasoning","dataset":"LingOly","model":"Llama 2 70B","rank_in_archive_order":11,"of":11,"metrics":{"Delta_NoContext":"1.1%","Exact Match Accuracy":"6.4%"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/2406.06196","atlas_url":"https://app.syntology.ai/?focus=2406.06196","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2406.06196"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/am-bean/lingOly","reach":null}],"summary":{"ran_violates":1,"ran_draft_wrong":1},"by_repo_kind":{"official":{"samples":2,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":2,"samples":[{"code_sha256_prefix":"0777d1ff44e7a37b","entry":"listtostr","repo":"am-bean/lingOly","repo_kind":"official","path":"testing/code/scoring.py","file_url":"https://github.com/am-bean/lingOly/blob/HEAD/testing/code/scoring.py","link_basis":"first_harvest_node","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"0777d1ff44e7a37b"}},{"code_sha256_prefix":"35d798e826fc7898","entry":"make_batch","repo":"am-bean/lingOly","repo_kind":"official","path":"testing/code/benchmark_model.py","file_url":"https://github.com/am-bean/lingOly/blob/HEAD/testing/code/benchmark_model.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"35d798e826fc7898"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}