Papers › Corpus Selection Approaches for Multilingual Parsing from Raw Text to Universal Dependencies

Corpus Selection Approaches for Multilingual Parsing from Raw Text to Universal Dependencies

1 Aug 2017CONLL 2017 8archive 2025-07-28

Ryan Hornby, Clark Taylor, Jungyeul Park

This paper describes UALing{'}s approach to the \textit{CoNLL 2017 UD Shared Task} using corpus selection techniques to reduce training data size. The methodology is simple: we use similarity measures to select a corpus from available training data (even from multiple corpora for surprise languages) and use the resulting corpus to complete the parsing task. The training and parsing is done with the baseline UDPipe system (Straka et al., 2016). While our approach reduces the size of training data significantly, it retains performance within 0.5{\%} of the baseline system. Due to the reduction in training data size, our system performs faster than the na{\"\i}ve, complete corpus method. Specifically, our system runs in less than 10 minutes, ranking it among the fastest entries for this task. Our system is available at \url{https://github.com/CoNLL-UD-2017/UALING}.

PaperPDFCode

Code

CoNLL-UD-2017/UALING officialmentioned in paper report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections