Papers › CogALex 2.0: Impact of Data Quality on Lexical-Semantic Relation Prediction

CogALex 2.0: Impact of Data Quality on Lexical-Semantic Relation Prediction

14 Dec 2021NeurIPS Data-Centric AI Workshop 2021 12archive 2025-07-28

Christian Lang, Lennart Wachowiak, Barbara Heinisch, Dagmar Gromann

Predicting lexical-semantic relations between word pairs has successfully been accomplished by pre-trained neural language models. An XLM-RoBERTa-based approach, for instance, achieved the best performance differentiating between hypernymy, synonymy, antonymy, and random relations in four languages in the CogALex-VI 2020 shared task. However, the results also revealed strong performance divergences between languages and confusions of specific relations, especially hypernymy and synonymy. Upon inspection, a difference in data quality across languages and relations could be observed. Thus, we provide a manually improved dataset for lexical-semantic relation prediction and evaluate its impact across three pre-trained neural language models.

PaperPDFCode

Code

Text2TCS/CogALex-2.0 mentioned in paper report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Hypernym DiscoveryRelation ClassificationRelation Prediction

1 archive task tag without a task page not shown.

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections