Datasets › GePaDe
GePaDe
This dataset encompasses 265 speeches (over 200,000 tokens) from the German Bundestag, primarily from the 19th legislative term (2017-2021), given by 195 distinct speakers representing 6 political parties.
The data was annotated to perform a semantic role labeling task, namely to identify who said what to whom (speaker attribution). Cues (triggers) were annotated that are associated with events of speech, writing, or thought. Additionally, the arguments (roles) of each trigger have been annotated, encompassing the SOURCE, ADDRESSEE, MESSAGE, MEDIUM, TOPIC, and EVIDENCE related to the speech event.
The dataset was introduced in the international GermEval 2023 Shared Task on Speaker Attribution in Newswire and Parliamentary Debates (SpkAtt-2023) to evaluate the quality of systems for automated identification of cues and associated roles.
Reference
Rehbein, I. et al, Overview of the GermEval 2023 Shared Task on Speaker Attribution in Newswire and Parliamentary Debates, https://github.com/umanlp/SpkAtt-2023/blob/master/doc/SpkAtt2023-proceedings.pdf
Benchmarks archive 2025-07-28
All 2 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Speaker Attribution in German Parliamentary Debates (GermEval 2023, subtask 1) | GePaDe | Llama 2 70 B QLoRa adapted F1 0.813 | Speaker attribution in German parliamentary debates with... | umanlp/spkatt-2023 +1 | 1 | Compare |
| Speaker Attribution in German Parliamentary Debates (GermEval 2023, subtask 2) | GePaDe | Llama 2 70 B QLoRa adapted F1 0.891 | Speaker attribution in German parliamentary debates with... | umanlp/spkatt-2023 +1 | 1 | Compare |
Papers archive 2025-07-28
1 shown of 1 paper with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 1. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| Speaker attribution in German parliamentary debates with QLoRA-adapted large language models | 2 | 2 | 18 Sep 2023 | not harvested |
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
License archive 2025-07-28
No licence recorded in the archive. Absence here is not a statement about the dataset's terms.
Modalities archive 2025-07-28
Languages archive 2025-07-28
Variants archive 2025-07-28
- GePaDe
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections