Papers › Szeged Corpus 2.5: Morphological Modifications in a Manually POS-tagged Hungarian Corpus

Szeged Corpus 2.5: Morphological Modifications in a Manually POS-tagged Hungarian Corpus

1 May 2014LREC 2014 5archive 2025-07-28

Veronika Vincze, Viktor Varga, Katalin Ilona Simk{\'o}, J{\'a}nos Zsibrita, {\'A}goston Nagy, Rich{\'a}rd Farkas, J{\'a}nos Csirik

The Szeged Corpus is the largest manually annotated database containing the possible morphological analyses and lemmas for each word form. In this work, we present its latest version, Szeged Corpus 2.5, in which the new harmonized morphological coding system of Hungarian has been employed and, on the other hand, the majority of misspelled words have been corrected and tagged with the proper morphological code. New morphological codes are introduced for participles, causative / modal / frequentative verbs, adverbial pronouns and punctuation marks, moreover, the distinction between common and proper nouns is eliminated. We also report some statistical data on the frequency of the new morphological codes. The new version of the corpus made it possible to train magyarlanc, a data-driven POS-tagger of Hungarian on a dataset with the new harmonized codes. According to the results, magyarlanc is able to achieve a state-of-the-art accuracy score on the 2.5 version as well.

PaperPDF

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

POS

Datasets

Introduced by this paper, per the archive.

Szeged Corpus

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections