Datasets › RuTermEval (Track 2)

RuTermEval (Track 2) (CL-RuTerm3)

23 Apr 2025 archive 2025-07-28

CL-RuTerm3 dataset is a novel resource featuring nested term annotations across six domains (the main one is computational linguistics, also mathematics, medicine, economics, literature studies, and agrochemistry), and the RuTermEval-2024 competition, designed to evaluate term extraction systems on this data. The CL-RuTerm3 dataset, comprising 1270 abstracts and 15 full-text articles (over 165k tokens with over 37k annotated entities), is the largest of its kind for Russian scientific texts. Terms are classified into three categories based on lexical and domain specificity: specific terms, common terms, and nomens. The dataset’s unique features, such as nested term markup and cross-domain coverage, enable more realistic evaluation of ATE systems.

Second track is devoted to Nested term extraction (in sequence labeling format) and classification (labels are specific, common, nomen).

Benchmarks archive 2025-07-28

All 2 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Nested Term Extraction RuTermEval (Track 2) full nested Scoreboard Class-agnostic F1 0.78 Methods for Recognizing Nested Terms fulstock/Methods-for-Recognizing-Nested-Terms 1 Compare
Nested Term Recognition from Flat Supervision RuTermEval (Track 2) lemm. inc. + early dmg Scoreboard Class-agnostic F1 0.7337 Methods for Recognizing Nested Terms fulstock/Methods-for-Recognizing-Nested-Terms 1 Compare

Papers archive 2025-07-28

1 shown of 1 paper with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 1. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Methods for Recognizing Nested Terms 1 2 22 Apr 2025 not harvested

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

No modality tagged.

Languages archive 2025-07-28

Variants archive 2025-07-28

  • RuTermEval (Track 2)

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections