{"url":"/task/document-level-closed-information-extraction","name":"Document-level Closed Information Extraction","slug":"document-level-closed-information-extraction","description_markdown":"Document-level closed information extraction (DocIE) is a subtask of information extraction that seeks to extract a set of triplets, or facts, of the form `(subject, relation, object)` from unstructured texts that are fully linked to a reference knowledge base, i.e., consistent with a predefined set of entities and relations from a knowledge base. DocIE entails tasks such as mention detection, entity typing, named entity recognition, entity disambiguation, entity linking, coreference resolution, and document-level relation extraction. DocIE is more challenging than sentence-level closed information extraction as it involves capturing long-range dependencies effectively to extract relations between entities that are further apart from each other in the text. Another difference is that DocIE necessitates a coreference resolution stage to group all the different mentions in the document referring to the same entity. DocIE is crucial for applications such as knowledge graph construction, question answering, knowledge discovery, or text summarization.\r\n\r\n<span class=\"description-source\">Source: [REXEL: An End-to-end Model for Document-Level Relation Extraction and Entity Linking](https://arxiv.org/abs/2404.12788)</span>","categories":[{"name":"Medical","url":"/area/medical"},{"name":"Natural Language Processing","url":"/area/natural-language-processing"}],"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","slug_source":"archive_url"},"counts":{"papers_tagged":2,"papers_with_code":2,"benchmarks":3,"benchmark_tables_in_archive":3,"benchmark_tables_shown":3,"benchmark_tables_withheld_as_spam":0,"benchmark_definition":"a leaderboard table with at least one row; benchmark_tables_shown also counts the zero-row tables; benchmark_tables_in_archive adds the tables withheld as spam","datasets":3,"subtasks":0,"parent_tasks":1},"benchmarks":[{"leaderboard":"/sota/document-level-closed-information-extraction","slug":"document-level-closed-information-extraction","dataset":"DocRED","dataset_url":"/dataset/docred","rows_in_archive":1,"metrics":["Relation F1"],"first_row_in_archive_order":{"model":"REXEL","paper_title":"REXEL: An End-to-end Model for Document-Level Relation Extraction and Entity Linking","paper_url":"/paper/rexel-an-end-to-end-model-for-document-level","paper_date":"2024-04-19","arxiv_id":"2404.12788","code_links":[{"title":"amazon-science/e2e-docie","url":"https://github.com/amazon-science/e2e-docie"}],"syntology":{"n":3,"n_ran":3,"n_unverified":0,"n_pointer_only":0}}},{"leaderboard":"/sota/document-level-closed-information-extraction-1","slug":"document-level-closed-information-extraction-1","dataset":"DWIE","dataset_url":"/dataset/dwie","rows_in_archive":1,"metrics":["F1-Hard"],"first_row_in_archive_order":{"model":"REXEL","paper_title":"REXEL: An End-to-end Model for Document-Level Relation Extraction and Entity Linking","paper_url":"/paper/rexel-an-end-to-end-model-for-document-level","paper_date":"2024-04-19","arxiv_id":"2404.12788","code_links":[{"title":"amazon-science/e2e-docie","url":"https://github.com/amazon-science/e2e-docie"}],"syntology":{"n":3,"n_ran":3,"n_unverified":0,"n_pointer_only":0}}},{"leaderboard":"/sota/document-level-closed-information-extraction-2","slug":"document-level-closed-information-extraction-2","dataset":"DocRED-IE","dataset_url":"/dataset/docred-ie","rows_in_archive":1,"metrics":["Relation F1"],"first_row_in_archive_order":{"model":"REXEL","paper_title":"REXEL: An End-to-end Model for Document-Level Relation Extraction and Entity Linking","paper_url":"/paper/rexel-an-end-to-end-model-for-document-level","paper_date":"2024-04-19","arxiv_id":"2404.12788","code_links":[{"title":"amazon-science/e2e-docie","url":"https://github.com/amazon-science/e2e-docie"}],"syntology":{"n":3,"n_ran":3,"n_unverified":0,"n_pointer_only":0}}}],"datasets":[{"url":"/dataset/docred","name":"DocRED","full_name":"","num_papers_in_archive":155},{"url":"/dataset/dwie","name":"DWIE","full_name":"Deutsche Welle corpus for Information Extraction","num_papers_in_archive":18},{"url":"/dataset/docred-ie","name":"DocRED-IE","full_name":"","num_papers_in_archive":1}],"subtasks":[],"parent_tasks":[{"url":"/task/information-extraction","name":"Information Extraction"}],"papers":{"order":"repositories listed in the archive (desc), then date (desc); the archive holds no stars","population":"papers tagged with this task that list at least one repository in the archive","shown":2,"of":2,"tagged_in_all":2,"items":[{"url":"/paper/2408-00103","title":"ReLiK: Retrieve and LinK, Fast and Accurate Entity Linking and Relation Extraction on an Academic Budget","date":"2024-07-31","arxiv_id":"2408.00103","repositories_listed":2,"syntology":null},{"url":"/paper/rexel-an-end-to-end-model-for-document-level","title":"REXEL: An End-to-end Model for Document-Level Relation Extraction and Entity Linking","date":"2024-04-19","arxiv_id":"2404.12788","repositories_listed":1,"syntology":{"n":3,"n_ran":3,"n_unverified":0,"n_pointer_only":0}}],"syntology_records":1,"syntology_note":"a paper without a record is not a recorded non-run: it may lack an arXiv id or simply be absent from the graph layer"},"description_links":{"kept":0,"unwrapped_to_text":0,"bare_urls_linked":0,"relative_images_dropped":0,"rule":"internal links are kept only when the target slug exists in the catalog"},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per-sample execution status on synthesized fixtures ('ran N of M samples'); not a correctness claim and not a ranking signal.","status_vocabulary":{"ran_honours":"ran, honoured the contract we drafted","ran_violates":"ran, violated the contract we drafted","ran_draft_wrong":"ran; our contract draft was wrong, not the code","ran_fixture":"ran; our fixture could not drive it","ran":"ran on a synthesized input","unverified":"unverified (harvested, no recorded run)"}},"not_shown":{"libraries":"the archive has no per-task library table","trend_sparklines":"the Trend column of the benchmarks table was a rendered image; it is not in the archive","social_and_latest_sorts":"stars and social signals are not in the archive"}}