{"url":"/dataset/sturm","name":"Sturm","full_name":null,"description_markdown":"# Digital Edition: Sturm Edition \r\n\r\nSource:\r\nSchrade, Torsten: „Startseite“, in: DER STURM. Digitale Quellenedition zur Geschichte der internationalen Avantgarde, erarbeitet und herausgegeben von Marjam Trautmann und Torsten Schrade. Mainz, Akademie der Wissenschaften und der Literatur, Version 1 vom 16. Jul. 2018.\r\n\r\nThis NER Dataset is available under the license\r\n[Creative Commons Attribution 4.0 International (CC-BY 4.0)](https://creativecommons.org/licenses/by/4.0/)\r\n\r\nFrom the Sturm Edition we have built a NER dataset. These are 174 letters from the years 1914-1922, which are available online in TEI format (see https://sturm-edition.de/id/S.0000001). It contains as tagged entities only persons, places and dates. From the original TEI files we build an NER dataset with tags distributed as shown in the following Table\r\n\r\nTag | # All | # Train | # Test | # Devel \r\n----|-------|---------|--------|--------\r\npers         | 930  | 763 | 83 | 84 \r\ndate         | 722 | 612 | 59 | 51  \r\nplace        | 492 | 374 | 59 | 59   \r\nnot tagged   | 33,809 | 27,047 | 3,306 | 3,456\r\n\r\nWe provide the dataset in two formats together with a partition into a train, dev, and testset. The first one is an easy format similar to the well-known CONLL-X format and the second one is an easy json format with the following structure:\r\n\r\nIt consists of a list of samples. Each sample is in turn a list of words or special characters. These in turn are represented as a two-element list, where the first element is the word itself and the second element is the corresponding target tag. Here is an example:\r\n\r\n[[['Peter','B-pers'],[Müller,'I-pers'],['lebt','O'],['in','O'], ['Frankfurt','B-place'],['am','I-place'],['Main','I-place'],['.','O']],[['Gebürtig','O'],['stammt','O'],['er','O'],['aus','O'],['Berlin','B-place']]","description_withheld":null,"homepage":"https://github.com/NEISSproject/NERDatasets/tree/main/Sturm","introduced_date":"2021-04-23","introduced_date_note":null,"introduced_by":{"paper":"/paper/optimizing-small-berts-trained-for-german-ner","title":"Optimizing small BERTs trained for German NER","first_author":"Jochen Zöllner","url":null},"license":{"name":"CC-BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Named Entity Recognition (NER)","url":"/task/named-entity-recognition-ner","datasets_with_task":"/datasets/task/named-entity-recognition-ner"}],"languages":[{"name":"German","url":"/datasets/language/german"}],"variants":["Sturm"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}