Datasets › LegalCore

LegalCore

Introduced by Kangda Wei et al. in LegalCore: A Dataset for Legal Documents Event Coreference Resolution18 Feb 2025 archive 2025-07-28

Recognizing events and their coreferential men- tions in a document is essential for understand- ing semantic meanings of text. The existing re- search on event coreference resolution is mostly limited to news articles. In this paper, we present the first dataset for the legal domain, LegalCore, which has been annotated with comprehensive event and event coreference in- formation. The legal contract documents we an- notated in this dataset are several times longer than news articles, with an average length of around 25k tokens per document. The anno- tations show that legal documents have dense event mentions and feature both short-distance and super long-distance coreference links be- tween event mentions. We further benchmark mainstream Large Language Models (LLMs) on this dataset for both event identification and event coreference resolution tasks, and find that this dataset poses significant challenges for both open-source and proprietary LLMs, which all perform significantly worse than a su- pervised baseline.

Benchmarks archive 2025-07-28

No leaderboard in the archive resolves to this dataset.

Papers archive 2025-07-28

No paper in the archive has a leaderboard row on this dataset; the archive counts 1 paper for it but never published that list.

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • LegalCore

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections