Papers › The merits of Universal Language Model Fine-tuning for Small Datasets -- a case with...

The merits of Universal Language Model Fine-tuning for Small Datasets -- a case with Dutch book reviews

2 Oct 2019arXiv:1910.00896archive 2025-07-28

Benjamin van der Burgh, Suzan Verberne

We evaluated the effectiveness of using language models, that were pre-trained in one domain, as the basis for a classification model in another domain: Dutch book reviews. Pre-trained language models have opened up new possibilities for classification tasks with limited labelled data, because representation can be learned in an unsupervised fashion. In our experiments we have studied the effects of training set size (100-1600 items) on the prediction accuracy of a ULMFiT classifier, based on a language models that we pre-trained on the Dutch Wikipedia. We also compared ULMFiT to Support Vector Machines, which is traditionally considered suitable for small collections. We found that ULMFiT outperforms SVM for all training set sizes and that satisfactory results (~90%) can be achieved using training sets that can be manually annotated within a few hours. We deliver both our new benchmark collection of Dutch book reviews for sentiment classification as well as the pre-trained Dutch language model to the community.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

benjaminvdb/110kDBRD mentioned on GitHub report
benjaminvdb/DBRD mentioned on GitHub report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

ClassificationGeneral ClassificationLanguage ModelingLanguage ModellingSentiment AnalysisSentiment Classification

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Methods

AWD-LSTMActivation RegularizationDiscriminative Fine-TuningDropConnectDropoutEmbedding DropoutLSTMSVMSigmoid ActivationSlanted Triangular Learning RatesTanh ActivationTemporal Activation RegularizationULMFiTVariational DropoutWeight Tying

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections