Home › Datasets › language › German

German datasets

archive 2025-07-28

199 datasets carry the language tag "German", ordered by the archive's paper count. Page 3 of 5: 48 shown of 199. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

German datasets 97–144 of 199

Operating rooms (ORs) are complex, high-stakes environments requiring precise understanding of interactions among medical staff, tools, and equipment for enhancing surgical assistance, situational awareness, and patient safety.
3 papers · 3 benchmarks
MRS (Multilingual Reply Suggestion)
MRS, a multilingual reply suggestion dataset with ten languages.
3 papers · 0 benchmarks
Multilingual automatic speech recognition (ASR) in the medical domain serves as a foundational task for various downstream applications such as speech translation, spoken language understanding, and voice-activated assistants.
3 papers · 0 benchmarks
Patzig contains handwritten texts written in modern German.
3 papers · 0 benchmarks
QALD-9-Plus Dataset Description QALD-9-Plus is the dataset for Knowledge Graph Question Answering (KGQA) based on well-known QALD-9.
3 papers · 1 benchmark
RWTH-PHOENIX Handshapes dev set (RWTH-PHOENIX-Weather 2014 MS Handshapes dev set)
We manually labelled 3359 images from the RWTH-PHOENIX-Weather 2014 Development set.
3 papers · 1 benchmark
Schiller (Shiller)
Schiller contains handwritten texts written in modern German.
3 papers · 0 benchmarks
Schwerin contains handwritten texts written in medieval German.
3 papers · 0 benchmarks
AM2iCo (Adversarial and Multilingual Meaning in Context)
AM2iCo is a wide-coverage and carefully designed cross-lingual and multilingual evaluation set.
2 papers · 0 benchmarks
APE (Automatic Post-Editing)
APE is useful to evaluate Machine Translation automatic post-editing (APE), which is the task of improving the output of a blackbox MT system by automatically fixing its mistakes.
2 papers · 0 benchmarks
Named entities in Bavarian text Details: Siyao Peng, Zihang Sun, Huangyan Shan, Marie Kolm, Verena Blaschke, Ekaterina Artemova, and Barbara Plank.
2 papers · 0 benchmarks
BioVid (BioVid Heat Pain Database)
To advance methods for pain assessment, in particular automatic assessment methods, the BioVid Heat Pain Database was collected in a collaboration of the Neuro-Information Technology group of the University of Magdeburg and the Medical…
2 papers · 0 benchmarks
DEplain-APA-sent: A German Parallel Corpus for Sentence Simplification on News Texts DEplain is a new dataset of parallel, professionally written and manually aligned simplifications in plain German “plain DE” (or in German: “Einfache…
2 papers · 1 benchmark
DEplain-web-sent: A German Parallel Corpus for Sentence Simplification on Web Texts DEplain is a new dataset of parallel, professionally written and manually aligned simplifications in plain German “plain DE” (or in German: “Einfache…
2 papers · 1 benchmark
Dataset of Legal Documents consists of court decisions from 2017 and 2018 were selected for the dataset, published online by the Federal Ministry of Justice and Consumer Protection.
2 papers · 0 benchmarks
DeCOCO is a bilingual (English-German) corpus of image descriptions, where the English part is extracted from the COCO dataset, and the German part are translations by a native German speaker.
2 papers · 0 benchmarks
The GermEval dataset is a valuable resource for natural language processing (NLP) tasks, specifically named entity recognition (NER), conducted in the German language.
2 papers · 0 benchmarks
GermanDPR is a dataset for passage retrieval in German.
2 papers · 0 benchmarks
InVar-100 (Industrial Objects in Varied Contexts)
The Industrial Objects in Varied Contexts (InVar) Dataset was internally produced by our team and contains 100 objects in 20800 total images (208 images per class).
2 papers · 0 benchmarks
Jam-ALT (JamALT: A Formatting-Aware Lyrics Transcription Benchmark)
JamALT is a revision of the JamendoLyrics dataset (80 songs in 4 languages), adapted for use as an automatic lyrics transcription (ALT) benchmark.
2 papers · 5 benchmarks
Kaleidoscope (Kaleidoscope: In-language Exams for Massively Multilingual Vision Evaluation)
The evaluation of vision-language models (VLMs) has mainly relied on English-language benchmarks, leaving significant gaps in both multilingual and multicultural coverage.
2 papers · 0 benchmarks
M2QA (Multi-domain Multilingual Question Answering)
M2QA (Multi-domain Multilingual Question Answering) is an extractive question answering benchmark for evaluating joint language and domain transfer.
2 papers · 0 benchmarks
MobIE is a German-language dataset which is human-annotated with 20 coarse- and fine-grained entity types and entity linking information for geographically linkable entities.
2 papers · 0 benchmarks
Morph Call is a suite of 46 probing tasks for four Indo-European languages that fall under different morphology: Russian, French, English, and German.
2 papers · 0 benchmarks
MuCo-VQA consist of large-scale (3.7M) multilingual and code-mixed VQA datasets in multiple languages: Hindi (hi), Bengali (bn), Spanish (es), German (de), French (fr) and code-mixed language pairs: en-hi, en-bn, en-fr, en-de and en-es.
2 papers · 0 benchmarks
MultiTACRED is a multilingual version of the large-scale TAC Relation Extraction Dataset.
2 papers · 0 benchmarks
PCC (Potsdam Commentary Corpus)
The Potsdam Commentary Corpus (PCC) is a corpus of 220 German newspaper commentaries (2.900 sentences, 44.000 tokens) taken from the online issues of the Märkische Allgemeine Zeitung (MAZ subcorpus) and Tagesspiegel (ProCon subcorpus) and…
2 papers · 0 benchmarks
A human-revised dataset for seven languages that allows for the evaluation of multilingual RE systems.
2 papers · 0 benchmarks
SubEdits is a human-annnoated post-editing dataset of neural machine translation outputs, compiled from in-house NMT outputs and human post-edits of subtitles form Rakuten Viki.
2 papers · 0 benchmarks
Tilde MODEL Corpus (Tilde Multilingual Open Data for European Languages)
Tilde MODEL Corpus is a multilingual corpora for European languages – particularly focused on the smaller languages.
2 papers · 0 benchmarks
Unbalance Classification Using Vibration Data (Vibration Measurements on a Rotating Shaft at Different Unbalance Strengths)
This dataset contains vibration data recorded on a rotating drive train.
2 papers · 0 benchmarks
WikiCaps is a large-scale multilingual but non-parallel data set for multimodal machine translation and retrieval.
2 papers · 0 benchmarks
X-WikiRE is a new, large-scale multilingual relation extraction dataset in which relation extraction is framed as a problem of reading comprehension to allow for generalization to unseen relations.
2 papers · 0 benchmarks
AGB-DE is a legal NLP corpus for the automated detection of potentially void clauses in German standard form consumer contracts.
1 paper · 1 benchmark
AQL-22 (Archive Query Log)
The Archive Query Log (AQL) is a previously unused, comprehensive query log collected at the Internet Archive over the last 25 years.
1 paper · 0 benchmarks
This paper analyses two hitherto unstudied sites sharing state-backed disinformation, Reliable Recent News (rrn.world) and WarOnFakes (waronfakes.com), which publish content in Arabic, Chinese, English, French, German, and Spanish.
1 paper · 0 benchmarks
Digital Edition: Essays from Hannah Arendt We have created a NER dataset from the digital edition "Sechs Essays" by Hannah Arendt.
1 paper · 0 benchmarks
Dataset Overview: 998 images and 4,208 annotations focusing on interaction with in-vehicle infotainment (IVI) systems.
1 paper · 0 benchmarks
CGHD1152 (Circuit Graph Hand Drawn 1152)
- 1152 Images - 144 Circuits - 12 Drafter - 48,563 Object (Symbol, Structural, Text) Annotations
1 paper · 0 benchmarks
CISOL (Construction Industry Steel Ordering Lists Dataset)
The Construction Industry Steel Ordering Lists (CISOL) dataset comprises table-centric, real-world documents from the construction industry, annotated to facilitate the testing and training of deep learning models for table detection (TD)…
1 paper · 2 benchmarks
The CareerCoach 2022 gold standard is available for download in the NIF and JSON format, and draws upon documents from a corpus of over 99,000 education courses which have been retrieved from 488 different education providers.
1 paper · 0 benchmarks
DEplain-APA-doc: A German Parallel Corpus for Document Simplification on News Texts DEplain is a new dataset of parallel, professionally written and manually aligned simplifications in plain German “plain DE” (or in German: “Einfache…
1 paper · 1 benchmark
DEplain-web-doc: A German Parallel Corpus for Document Simplification on Web Texts DEplain is a new dataset of parallel, professionally written and manually aligned simplifications in plain German “plain DE” (or in German: “Einfache…
1 paper · 1 benchmark
Dubbing Test Set consists of two subsets extracted from the En→De test set of COVOST-2, a large-scale multilingual speech translation corpus based on Common Voice.
1 paper · 0 benchmarks
ELMTEX Dataset (ELMTEX Dataset: Fine-Tuning Large Language Models for Structured Clinical Information Extraction)
We introduced a new dataset of clinical report summaries, annotated with structured information across 15 categories.
1 paper · 0 benchmarks
Dataset Description EUROPA is a dataset designed for training and evaluating multilingual keyphrase generation models in the legal domain.
1 paper · 0 benchmarks
Predictions of energy consumption are crucial for energy retailers to minimize deviations from energy acquired in the day-ahead market and the actual consumption of their customers.
1 paper · 0 benchmarks
Contains 1000 semantic queries and the corresponding English, German and Portuguese verbalizations for EventKG - an event-centric knowledge graph with more than 970 thousand events.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.