{"url":"/dataset/codred","name":"CodRED","full_name":null,"description_markdown":"CodRED is the first human-annotated cross-document relation extraction (RE) dataset, aiming to test the RE systems’ ability of knowledge acquisition in the wild. CodRED has the following features:\r\n\r\n* it requires natural language understanding in different granularity, including coarse-grained document retrieval, as well as fine-grained cross-document multi-hop reasoning;\r\n\r\n* it contains 30,504 relational facts associated with 210,812 reasoning text paths, as well as enjoys a broad range of balanced relations, and long documents in diverse topics;\r\n\r\n* it provides strong supervision about the reasoning text paths for predicting the relation, to help guide RE systems to perform meaningful and interpretable reasoning;\r\n* it contains adversarially-created hard NA instances to avoid RE models to predict relations by inferring from entity names instead of text information.","description_withheld":null,"homepage":"https://github.com/thunlp/CodRED","introduced_date":"2021-11-01","introduced_date_note":null,"introduced_by":{"paper":"/paper/codred-a-cross-document-relation-extraction","title":"CodRED: A Cross-Document Relation Extraction Dataset for Acquiring Knowledge in the Wild","first_author":"Yuan YAO","url":null},"license":{"name":"MIT","url":null},"modalities":[],"tasks":[],"languages":[],"variants":["CodRED"],"data_loaders":[],"num_papers_in_archive":5,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}