{"url":"/dataset/szeged-corpus","name":"Szeged Corpus","full_name":null,"description_markdown":"The Szeged Treebank is the largest fully manually annotated treebank of the Hungarian language. It contains 82,000 sentences, 1.2 million words and 250,000 punctuation marks. Texts were selected from six different domains, ~200,000 words in size from each. The domains are the following:\r\n\r\nfiction\r\ncompositions of pupils between 14-16 years of age\r\nnewspaper articles (from the newspapers Népszabadság, Népszava, Magyar Hírlap, HVG)\r\ntexts in informatics\r\nlegal texts\r\nbusiness and financial news\r\nThe treebank exists in three versions:\r\n\r\nSzeged Treebank 1.0 is annotated for noun phrases and clauses;\r\nSzeged Treebank 2.0 contains a deep phrase-structured syntactic analysis for all sentences;\r\nSzeged Dependency Treebank contains dependency-style annotation of all sentences.\r\nA morphologically reannotated version of the corpus, Szeged Corpus 2.5 has just been released, where participles, causative, frequentative and model verbs are distinctively marked, and unknown or misspelled words have been corrected, along with some minor morphological modifications.\r\nIf you are interested in Szeged Corpus 2.5, please contact Veronika Vincze.","description_withheld":null,"homepage":"https://rgai.inf.u-szeged.hu/node/113","introduced_date":"2004-08-23","introduced_date_note":null,"introduced_by":{"paper":"/paper/szeged-corpus-25-morphological-modifications","title":"Szeged Corpus 2.5: Morphological Modifications in a Manually POS-tagged Hungarian Corpus","first_author":"Veronika Vincze","url":null},"license":{"name":"https://rgai.inf.u-szeged.hu/node/113","url":"https://rgai.inf.u-szeged.hu/node/113"},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Semantic Parsing","url":"/task/semantic-parsing","datasets_with_task":"/datasets/task/semantic-parsing"},{"name":"Part-Of-Speech Tagging","url":"/task/part-of-speech-tagging","datasets_with_task":"/datasets/task/part-of-speech-tagging"},{"name":"Morphological Tagging","url":"/task/morphological-tagging","datasets_with_task":"/datasets/task/morphological-tagging"},{"name":"Lemmatization","url":"/task/lemmatization","datasets_with_task":"/datasets/task/lemmatization"}],"languages":[{"name":"Hungarian","url":"/datasets/language/hungarian"}],"variants":["Szeged Corpus"],"data_loaders":[],"num_papers_in_archive":4,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}