{"url":"/dataset/multitacred","name":"MultiTACRED","full_name":null,"description_markdown":"MultiTACRED is a multilingual version of the large-scale \r\n[TAC Relation Extraction Dataset](https://nlp.stanford.edu/projects/tacred). It covers 12 typologically diverse \r\nlanguages from 9 language families, and was created by the \r\n[Speech & Language Technology group of DFKI](https://www.dfki.de/slt) by machine-translating the instances of the \r\noriginal TACRED dataset and automatically projecting their entity annotations. For details of the original TACRED's \r\ndata collection and annotation process, see the [Stanford paper](https://aclanthology.org/D17-1004/). Translations are \r\nsyntactically validated by checking the correctness of the XML tag markup. Any translations with an invalid tag \r\nstructure, e.g. missing or invalid head or tail tag pairs, are discarded (on average, 2.3% of the instances).\r\n\r\nLanguages covered are: Arabic, Chinese, Finnish, French, German, Hindi, Hungarian, Japanese, Polish,\r\n Russian, Spanish, Turkish. Intended use is supervised relation classification. Audience - researchers.\r\n\r\nMultiTACRED is released via the Linguistic Data Consortium (LDC License). You can download MultiTACRED from the [LDC MultiTACRED webpage](https://catalog.ldc.upenn.edu/LDC2024T09).\r\n\r\n Please see [our ACL paper](https://arxiv.org/abs/2305.04582) for full details.","description_withheld":null,"homepage":"https://github.com/DFKI-NLP/MultiTACRED","introduced_date":"2023-05-08","introduced_date_note":null,"introduced_by":{"paper":"/paper/multitacred-a-multilingual-version-of-the-tac","title":"MultiTACRED: A Multilingual Version of the TAC Relation Extraction Dataset","first_author":"Leonhard Hennig","url":null},"license":{"name":"LDC","url":"https://catalog.ldc.upenn.edu/license/ldc-non-members-agreement.pdf"},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Relation Extraction","url":"/task/relation-extraction","datasets_with_task":"/datasets/task/relation-extraction"},{"name":"Relation Classification","url":"/task/relation-classification","datasets_with_task":"/datasets/task/relation-classification"}],"languages":[{"name":"French","url":"/datasets/language/french"},{"name":"Spanish","url":"/datasets/language/spanish"},{"name":"German","url":"/datasets/language/german"},{"name":"Chinese","url":"/datasets/language/chinese"},{"name":"Japanese","url":"/datasets/language/japanese"},{"name":"Russian","url":"/datasets/language/russian"},{"name":"Arabic","url":"/datasets/language/arabic"},{"name":"Finnish","url":"/datasets/language/finnish"},{"name":"Hindi","url":"/datasets/language/hindi"},{"name":"Hungarian","url":"/datasets/language/hungarian"},{"name":"Polish","url":"/datasets/language/polish"},{"name":"Turkish","url":"/datasets/language/turkish"}],"variants":["MultiTACRED"],"data_loaders":[{"repo":"https://github.com/huggingface/datasets","url":"https://huggingface.co/datasets/DFKI-SLT/multitacred","frameworks":["tf","pytorch","jax"]}],"num_papers_in_archive":2,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-25T09:33:49+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}