{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/the-sigmorphon-2022-shared-task-on-morpheme","title":"The SIGMORPHON 2022 Shared Task on Morpheme Segmentation","arxiv_id":"2206.07615","date":"2022-06-15","proceeding":"NAACL (SIGMORPHON) 2022 7","authors":["Khuyagbaatar Batsuren","Gábor Bella","Aryaman Arora","Viktor Martinović","Kyle Gorman","Zdeněk Žabokrtský","Amarsanaa Ganbold","Šárka Dohnalová","Magda Ševčíková","Kateřina Pelegrinová","Fausto Giunchiglia","Ryan Cotterell","Ekaterina Vylomova"],"abstract":"The SIGMORPHON 2022 shared task on morpheme segmentation challenged systems to decompose a word into a sequence of morphemes and covered most types of morphology: compounds, derivations, and inflections. Subtask 1, word-level morpheme segmentation, covered 5 million words in 9 languages (Czech, English, Spanish, Hungarian, French, Italian, Russian, Latin, Mongolian) and received 13 system submissions from 7 teams and the best system averaged 97.29% F1 score across all languages, ranging English (93.84%) to Latin (99.38%). Subtask 2, sentence-level morpheme segmentation, covered 18,735 sentences in 3 languages (Czech, English, Mongolian), received 10 system submissions from 3 teams, and the best systems outperformed all three state-of-the-art subword tokenization methods (BPE, ULM, Morfessor2) by 30.71% absolute. To facilitate error analysis and support any type of future studies, we released all system predictions, the evaluation script, and all gold standard datasets.","url_abs":"https://arxiv.org/abs/2206.07615v1","url_pdf":"https://arxiv.org/pdf/2206.07615v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"the-sigmorphon-2022-shared-task-on-morpheme","repo_url":"https://github.com/sigmorphon/2022segmentationst","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"all","task_name":"All"},{"task_slug":"morpheme-segmentaiton","task_name":"Morpheme Segmentaiton"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"sentence","task_name":"Sentence"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/morpheme-segmentaiton-on-unimorph-4-0","task":"Morpheme Segmentaiton","dataset":"UniMorph 4.0","model":"Morfessor2","rank_in_archive_order":8,"of":19,"metrics":{"f1 macro avg (subtask 2)":"50.65","lev dist (subtask 2)":"12.08","macro avg (subtask 1)":"25.57"},"uses_additional_data":false},{"leaderboard":"/sota/morpheme-segmentaiton-on-unimorph-4-0","task":"Morpheme Segmentaiton","dataset":"UniMorph 4.0","model":"ULM","rank_in_archive_order":9,"of":19,"metrics":{"f1 macro avg (subtask 2)":"45.99","lev dist (subtask 2)":"14.28","macro avg (subtask 1)":"20.61"},"uses_additional_data":false},{"leaderboard":"/sota/morpheme-segmentaiton-on-unimorph-4-0","task":"Morpheme Segmentaiton","dataset":"UniMorph 4.0","model":"WordPiece","rank_in_archive_order":10,"of":19,"metrics":{"f1 macro avg (subtask 2)":"40.59","lev dist (subtask 2)":"17.54","macro avg (subtask 1)":"15.89"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/2206.07615","atlas_url":"https://app.syntology.ai/?focus=2206.07615","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}