Papers › Neural Morphology Dataset and Models for Multiple Languages, from the Large to the Endangered

Neural Morphology Dataset and Models for Multiple Languages, from the Large to the Endangered

26 May 2021NoDaLiDa 2021 5arXiv:2105.12428archive 2025-07-28

Mika Hämäläinen, Niko Partanen, Jack Rueter, Khalid Alnajjar

We train neural models for morphological analysis, generation and lemmatization for morphologically rich languages. We present a method for automatically extracting substantially large amount of training data from FSTs for 22 languages, out of which 17 are endangered. The neural models follow the same tagset as the FSTs in order to make it possible to use them as fallback systems together with the FSTs. The source code, models and datasets have been released on Zenodo.

PaperPDFConference PDFCode

Code

mikahama/uralicNLP officialmentioned in paper report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

LemmatizationMorphological Analysis

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections