Papers › Challenges in Annotating and Parsing Spoken, Code-switched, Frisian-Dutch Data

Challenges in Annotating and Parsing Spoken, Code-switched, Frisian-Dutch Data

1 Apr 2021EACL (AdaptNLP) 2021 4archive 2025-07-28

Anouck Braggaar, Rob van der Goot

While high performance have been obtained for high-resource languages, performance on low-resource languages lags behind. In this paper we focus on the parsing of the low-resource language Frisian. We use a sample of code-switched, spontaneously spoken data, which proves to be a challenging setup. We propose to train a parser specifically tailored towards the target domain, by selecting instances from multiple treebanks. Specifically, we use Latent Dirichlet Allocation (LDA), with word and character N-grams. We use a deep biaffine parser initialized with mBERT. The best single source treebank (nl_alpino) resulted in an LAS of 54.7 whereas our data selection outperformed the single best transfer treebank and led to 55.6 LAS on the test data. Additional experiments consisted of removing diacritics from our Frisian data, creating more similar training data by cropping sentences and running our best model using XLM-R. These experiments did not lead to a better performance.

PaperPDFCode

Code

anouck96/parsingfrisian officialmentioned in paper report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

XLM-R

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Methods

XLM-RmBERT

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections