{"url":"/task/lemmatization","name":"Lemmatization","slug":"lemmatization","description_markdown":"**Lemmatization** is a process of determining a base or dictionary form (lemma) for a given surface form. Especially for languages with rich morphology it is important to be able to normalize words into their base forms to better support for example search engines and linguistic studies. Main difficulties in Lemmatization arise from encountering previously unseen words during inference time as well as disambiguating ambiguous surface forms which can be inflected variants of several different base forms depending on the context.\n\n\n<span class=\"description-source\">Source: [Universal Lemmatizer: A Sequence to Sequence Model for Lemmatizing Universal Dependencies Treebanks ](https://arxiv.org/abs/1902.00972)</span>","categories":[{"name":"Natural Language Processing","url":"/area/natural-language-processing"}],"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","slug_source":"archive_url"},"counts":{"papers_tagged":351,"papers_with_code":68,"benchmarks":0,"benchmark_tables_in_archive":0,"benchmark_tables_shown":0,"benchmark_tables_withheld_as_spam":0,"benchmark_definition":"a leaderboard table with at least one row; benchmark_tables_shown also counts the zero-row tables; benchmark_tables_in_archive adds the tables withheld as spam","datasets":3,"subtasks":0,"parent_tasks":0},"benchmarks":[],"datasets":[{"url":"/dataset/gum","name":"GUM","full_name":"Georgetown University Multilayer corpus","num_papers_in_archive":13},{"url":"/dataset/amalgum","name":"AMALGUM","full_name":"A Machine Annotated Lookalike of GUM","num_papers_in_archive":5},{"url":"/dataset/szeged-corpus","name":"Szeged Corpus","full_name":"","num_papers_in_archive":4}],"subtasks":[],"parent_tasks":[],"papers":{"order":"repositories listed in the archive (desc), then date (desc); the archive holds no stars","population":"papers tagged with this task that list at least one repository in the archive","shown":30,"of":68,"tagged_in_all":351,"items":[{"url":"/paper/stanza-a-python-natural-language-processing","title":"Stanza: A Python Natural Language Processing Toolkit for Many Human Languages","date":"2020-03-16","arxiv_id":"2003.07082","repositories_listed":5,"syntology":{"n":30,"n_ran":16,"n_unverified":14,"n_pointer_only":28}},{"url":"/paper/advancing-hungarian-text-processing-with","title":"Advancing Hungarian Text Processing with HuSpaCy: Efficient and Accurate NLP Pipelines","date":"2023-08-24","arxiv_id":"2308.12635","repositories_listed":2,"syntology":null},{"url":"/paper/sentence-embedding-models-for-ancient-greek","title":"Sentence Embedding Models for Ancient Greek Using Multilingual Knowledge Distillation","date":"2023-08-24","arxiv_id":"2308.13116","repositories_listed":2,"syntology":null},{"url":"/paper/top2vec-distributed-representations-of-topics","title":"Top2Vec: Distributed Representations of Topics","date":"2020-08-19","arxiv_id":"2008.09470","repositories_listed":2,"syntology":{"n":2,"n_ran":2,"n_unverified":0,"n_pointer_only":0}},{"url":"/paper/improving-lemmatization-of-non-standard","title":"Improving Lemmatization of Non-Standard Languages with Joint Learning","date":"2019-03-16","arxiv_id":"1903.06939","repositories_listed":2,"syntology":null},{"url":"/paper/lemmatag-jointly-tagging-and-lemmatizing-for","title":"LemmaTag: Jointly Tagging and Lemmatizing for Morphologically-Rich Languages with BRNNs","date":"2018-08-10","arxiv_id":"1808.03703","repositories_listed":2,"syntology":null},{"url":"/paper/parsipy-nlp-toolkit-for-historical-persian","title":"ParsiPy: NLP Toolkit for Historical Persian Texts in Python","date":"2025-03-22","arxiv_id":"2503.17810","repositories_listed":1,"syntology":null},{"url":"/paper/a-state-of-the-art-morphosyntactic-parser-and","title":"A State-of-the-Art Morphosyntactic Parser and Lemmatizer for Ancient Greek","date":"2024-10-15","arxiv_id":"2410.12055","repositories_listed":1,"syntology":null},{"url":"/paper/2409-13920","title":"One Model is All You Need: ByT5-Sanskrit, a Unified Model for Sanskrit NLP Tasks","date":"2024-09-20","arxiv_id":"2409.13920","repositories_listed":1,"syntology":null},{"url":"/paper/open-source-web-service-with-morphological","title":"Open-Source Web Service with Morphological Dictionary-Supplemented Deep Learning for Morphosyntactic Analysis of Czech","date":"2024-06-18","arxiv_id":"2406.12422","repositories_listed":1,"syntology":null},{"url":"/paper/heidelberg-boston-sigtyp-2024-shared-task","title":"Heidelberg-Boston @ SIGTYP 2024 Shared Task: Enhancing Low-Resource Language Analysis With Character-Aware Hierarchical Transformers","date":"2024-05-30","arxiv_id":"2405.20145","repositories_listed":1,"syntology":null},{"url":"/paper/opera-graeca-adnotata-building-a-34m-token","title":"Opera Graeca Adnotata: Building a 34M+ Token Multilayer Corpus for Ancient Greek","date":"2024-03-31","arxiv_id":"2404.00739","repositories_listed":1,"syntology":null},{"url":"/paper/cross-lingual-named-entity-corpus-for-slavic","title":"Cross-lingual Named Entity Corpus for Slavic Languages","date":"2024-03-30","arxiv_id":"2404.00482","repositories_listed":1,"syntology":null},{"url":"/paper/evaluating-shortest-edit-script-methods-for","title":"Evaluating Shortest Edit Script Methods for Contextual Lemmatization","date":"2024-03-25","arxiv_id":"2403.16968","repositories_listed":1,"syntology":null},{"url":"/paper/banlemma-a-word-formation-dependent-rule-and","title":"BanLemma: A Word Formation Dependent Rule and Dictionary Based Bangla Lemmatizer","date":"2023-11-06","arxiv_id":"2311.03078","repositories_listed":1,"syntology":null},{"url":"/paper/lexicon-and-rule-based-word-lemmatization","title":"Lexicon and Rule-based Word Lemmatization Approach for the Somali Language","date":"2023-08-03","arxiv_id":"2308.01785","repositories_listed":1,"syntology":null},{"url":"/paper/hybrid-lemmatization-in-huspacy","title":"Hybrid lemmatization in HuSpaCy","date":"2023-06-13","arxiv_id":"2306.07636","repositories_listed":1,"syntology":null},{"url":"/paper/exploring-large-language-models-for-classical","title":"Exploring Large Language Models for Classical Philology","date":"2023-05-23","arxiv_id":"2305.13698","repositories_listed":1,"syntology":null},{"url":"/paper/brent-bidirectional-retrieval-enhanced","title":"BRENT: Bidirectional Retrieval Enhanced Norwegian Transformer","date":"2023-04-19","arxiv_id":"2304.09649","repositories_listed":1,"syntology":null},{"url":"/paper/exploring-the-state-of-the-art-language","title":"Transformers on Multilingual Clause-Level Morphology","date":"2022-11-03","arxiv_id":"2211.01736","repositories_listed":1,"syntology":null},{"url":"/paper/knowledge-authoring-with-factual-english","title":"Knowledge Authoring with Factual English","date":"2022-08-05","arxiv_id":"2208.03094","repositories_listed":1,"syntology":null},{"url":"/paper/dadmatools-natural-language-processing","title":"DadmaTools: Natural Language Processing Toolkit for Persian Language","date":"2022-07-01","arxiv_id":null,"repositories_listed":1,"syntology":null},{"url":"/paper/stylistic-fingerprints-pos-tags-and-inflected","title":"Stylistic Fingerprints, POS-tags and Inflected Languages: A Case Study in Polish","date":"2022-06-05","arxiv_id":"2206.02208","repositories_listed":1,"syntology":null},{"url":"/paper/huspacy-an-industrial-strength-hungarian","title":"HuSpaCy: an industrial-strength Hungarian natural language processing toolkit","date":"2022-01-06","arxiv_id":"2201.01956","repositories_listed":1,"syntology":{"n":2,"n_ran":0,"n_unverified":2,"n_pointer_only":0}},{"url":"/paper/elit-emory-language-and-information-toolkit","title":"ELIT: Emory Language and Information Toolkit","date":"2021-09-08","arxiv_id":"2109.03903","repositories_listed":1,"syntology":null},{"url":"/paper/lemmatization-of-historical-old-literary","title":"Lemmatization of Historical Old Literary Finnish Texts in Modern Orthography","date":"2021-07-07","arxiv_id":"2107.03266","repositories_listed":1,"syntology":null},{"url":"/paper/neural-morphology-dataset-and-models-for","title":"Neural Morphology Dataset and Models for Multiple Languages, from the Large to the Endangered","date":"2021-05-26","arxiv_id":"2105.12428","repositories_listed":1,"syntology":null},{"url":"/paper/enhancing-sequence-to-sequence-neural","title":"Enhancing Sequence-to-Sequence Neural Lemmatization with External Resources","date":"2021-01-28","arxiv_id":"2101.12056","repositories_listed":1,"syntology":null},{"url":"/paper/dbtagger-multi-task-learning-for-keyword","title":"DBTagger: Multi-Task Learning for Keyword Mapping in NLIDBs Using Bi-Directional Recurrent Neural Networks","date":"2021-01-11","arxiv_id":"2101.04226","repositories_listed":1,"syntology":null},{"url":"/paper/trankit-a-light-weight-transformer-based","title":"Trankit: A Light-Weight Transformer-based Toolkit for Multilingual Natural Language Processing","date":"2021-01-09","arxiv_id":"2101.03289","repositories_listed":1,"syntology":null}],"syntology_records":3,"syntology_note":"a paper without a record is not a recorded non-run: it may lack an arXiv id or simply be absent from the graph layer"},"description_links":{"kept":0,"unwrapped_to_text":0,"bare_urls_linked":0,"relative_images_dropped":0,"rule":"internal links are kept only when the target slug exists in the catalog"},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per-sample execution status on synthesized fixtures ('ran N of M samples'); not a correctness claim and not a ranking signal.","status_vocabulary":{"ran_honours":"ran, honoured the contract we drafted","ran_violates":"ran, violated the contract we drafted","ran_draft_wrong":"ran; our contract draft was wrong, not the code","ran_fixture":"ran; our fixture could not drive it","ran":"ran on a synthesized input","unverified":"unverified (harvested, no recorded run)"}},"not_shown":{"libraries":"the archive has no per-task library table","trend_sparklines":"the Trend column of the benchmarks table was a rendered image; it is not in the archive","social_and_latest_sorts":"stars and social signals are not in the archive"}}