Papers › BERTić -- The Transformer Language Model for Bosnian, Croatian, Montenegrin and Serbian

BERTić -- The Transformer Language Model for Bosnian, Croatian, Montenegrin and Serbian

19 Apr 2021arXiv:2104.09243archive 2025-07-28

Nikola Ljubešić, Davor Lauc

In this paper we describe a transformer model pre-trained on 8 billion tokens of crawled text from the Croatian, Bosnian, Serbian and Montenegrin web domains. We evaluate the transformer model on the tasks of part-of-speech tagging, named-entity-recognition, geo-location prediction and commonsense causal reasoning, showing improvements on all tasks over state-of-the-art models. For commonsense reasoning evaluation, we introduce COPA-HR -- a translation of the Choice of Plausible Alternatives (COPA) dataset into Croatian. The BERTi\'c model is made available for free usage and further task-specific fine-tuning through HuggingFace.

PaperPDF

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Commonsense Causal ReasoningLanguage ModelingLanguage ModellingNamed Entity RecognitionNamed Entity Recognition (NER)Part-Of-Speech TaggingTranslationnamed-entity-recognition

Datasets

Introduced by this paper, per the archive.

COPA-HR

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections