Papers › Contract Discovery: Dataset and a Few-Shot Semantic Retrieval Challenge with...

Contract Discovery: Dataset and a Few-Shot Semantic Retrieval Challenge with Competitive Baselines

10 Nov 2019Findings of the Association for Computational Linguistics 2020arXiv:1911.03911archive 2025-07-28

Łukasz Borchmann, Dawid Wiśniewski, Andrzej Gretkowski, Izabela Kosmala, Dawid Jurkiewicz, Łukasz Szałkiewicz, Gabriela Pałka, Karol Kaczmarek, Agnieszka Kaliska, Filip Graliński

We propose a new shared task of semantic retrieval from legal texts, in which a so-called contract discovery is to be performed, where legal clauses are extracted from documents, given a few examples of similar clauses from other legal acts. The task differs substantially from conventional NLI and shared tasks on legal information extraction (e.g., one has to identify text span instead of a single document, page, or paragraph). The specification of the proposed task is followed by an evaluation of multiple solutions within the unified framework proposed for this branch of methods. It is shown that state-of-the-art pretrained encoders fail to provide satisfactory results on the task proposed. In contrast, Language Model-based solutions perform better, especially when unsupervised fine-tuning is applied. Besides the ablation studies, we addressed questions regarding detection accuracy for relevant text fragments depending on the number of examples available. In addition to the dataset and reference results, LMs specialized in the legal domain were made publicly available.

PaperPDFConference PDFCode

Code

applicaai/contract-discovery officialmentioned in papermentioned on GitHub report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Few-Shot LearningLanguage ModelingLanguage ModellingRetrievalSemantic RetrievalSemantic Similarity

Datasets

Introduced by this paper, per the archive.

Contract Discovery

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Semantic Retrieval Contract Discovery Human baseline Soft-F1 0.84 #1 of 6 Archive leaderboard report
Semantic Retrieval Contract Discovery k-NN with sentence n-grams, GPT-2 embeddings, fICA Soft-F1 0.51 #2 of 6 Archive leaderboard report
Semantic Retrieval Contract Discovery LSA baseline Soft-F1 0.39 #4 of 6 Archive leaderboard report
Semantic Retrieval Contract Discovery Universal Sentence Encoder Soft-F1 0.38 #5 of 6 Archive leaderboard report
Semantic Retrieval Contract Discovery Sentence BERT Soft-F1 0.31 #6 of 6 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

k-NN

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections