{"url":"/dataset/germeval-2021-toxic-engaging-fact-claiming","name":"GermEval","full_name":null,"description_markdown":"The **GermEval dataset** is a valuable resource for natural language processing (NLP) tasks, specifically **named entity recognition (NER)**, conducted in the German language. Here are some key details about this dataset:\r\n\r\n- **Task**: **Token Classification** (specifically, named entity recognition)\r\n- **Language**: **German**\r\n- **Size**: The dataset falls within the category of **100K < n < 1M** tokens.\r\n- **Source**: The data was sampled from **German Wikipedia** and **News Corpora**, comprising a collection of citations.\r\n- **Annotations**: The annotations were created through **crowdsourcing** efforts.\r\n- **License**: The dataset is available under the **cc-by-4.0** license.\r\n- **Content**: It covers over **31,000 sentences**, corresponding to more than **590,000 tokens**.\r\n- **Purpose**: Researchers and practitioners can use this dataset to train and evaluate NER models for German text.\r\n\r\nYou can find more information and explore the dataset on the [Hugging Face Datasets page](https://huggingface.co/datasets/germeval_14) ¹.\r\n\r\n(1) germeval_14 · Datasets at Hugging Face. https://huggingface.co/datasets/germeval_14.\r\n(2) GermEval-2018 Corpus (DE) - Empirical Linguistics and ... - heiDATA. https://heidata.uni-heidelberg.de/dataset.xhtml?persistentId=doi:10.11588/data/0B5VML.\r\n(3) GermEval 2014 Named Entity Recognition Shared Task - Data and Task Setup. https://sites.google.com/site/germeval2014ner/data.\r\n(4) 6 Best German Language Datasets of 2022 | Twine - Twine Blog. https://www.twine.net/blog/best-german-language-datasets/.\r\n(5) germeval_14 | TensorFlow Datasets. https://www.tensorflow.org/datasets/community_catalog/huggingface/germeval_14.\r\n(6) undefined. http://www.stern.de/sport/fussball/krawalle-in-der-fussball-bundesliga-dfb-setzt-auf-falsche-konzepte-1553657.html.\r\n(7) undefined. http://www.fr-online.de/in_und_ausland/sport/aktuell/1618625_Frings-schaut-finster-in-die-Zukunft.html.","description_withheld":null,"homepage":"https://germeval.github.io/","introduced_date":null,"introduced_date_note":null,"introduced_by":null,"license":null,"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Classification of toxic, engaging, fact-claiming comments","url":"/task/classification-of-toxic-engaging-fact","datasets_with_task":"/datasets/task/classification-of-toxic-engaging-fact"}],"languages":[{"name":"German","url":"/datasets/language/german"}],"variants":["GermEval"],"data_loaders":[],"num_papers_in_archive":2,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}