{"url":"/dataset/gtc-topic-based-paragraph-classification","name":"Genocide Transcript Corpus (GTC): Topic-Based Paragraph Classification in Genocide-Related Court Transcripts","full_name":null,"description_markdown":"The **Topic-Based Paragraph Classification in Genocide-Related Court Transcripts (GTC)** dataset is the first reference corpus annotated with samples from genocide tribunals in different international criminal courts. It is made up of witness statements about violence experienced. The material consists of 1475 text passages with about 40 to 120 pages per transcript, covering 3 tribunals: the Extraordinary Chambers in the Courts of Cambodia (ECCC) - 438 pages, the International Criminal Tribunal for Rwanda (ICTR) - 566 pages, and the International Criminal Tribunal of the Former Yugoslavia (ICTY) - 416 pages. As no datasets of any kind containing genocide court transcripts have been published nor other forms of pre-structured or annotated text data in this field of research exist, the aim was to address this gap by providing a systematically annotated dataset. \r\n\r\nPotential use cases include genocide-related inquiry conducted by those who need better to access, explore, and search through extensive documentation on these topics including researchers, lawyers and other practitioners. Broadly, its stated aim is to serve 3 purposes:\r\n\r\n* (1) to provide a first reference corpus for the community\r\n* (2) to establish benchmark performances (using state-of-the-art transformer-based approaches) for the new classification task of paragraph identification of violence-related\r\nwitness statements\r\n* (3) to explore first steps towards transfer learning within the domain","description_withheld":null,"homepage":"https://github.com/miriamschirmer/genocide-transcript-corpus","introduced_date":"2022-04-06","introduced_date_note":null,"introduced_by":{"paper":"/paper/a-new-dataset-for-topic-based-paragraph","title":"A New Dataset for Topic-Based Paragraph Classification in Genocide-Related Court Transcripts","first_author":"Miriam Schirmer","url":null},"license":null,"modalities":[],"tasks":[{"name":"Text Classification","url":"/task/text-classification","datasets_with_task":"/datasets/task/text-classification"}],"languages":[],"variants":["Genocide Transcript Corpus (GTC): Topic-Based Paragraph Classification in Genocide-Related Court Transcripts"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}