Papers › Czech Dataset for Cross-lingual Subjectivity Classification

Czech Dataset for Cross-lingual Subjectivity Classification

29 Apr 2022LREC 2022 6arXiv:2204.13915archive 2025-07-28

Pavel Přibáň, Josef Steinberger

In this paper, we introduce a new Czech subjectivity dataset of 10k manually annotated subjective and objective sentences from movie reviews and descriptions. Our prime motivation is to provide a reliable dataset that can be used with the existing English dataset as a benchmark to test the ability of pre-trained multilingual models to transfer knowledge between Czech and English and vice versa. Two annotators annotated the dataset reaching 0.83 of the Cohen's \k{appa} inter-annotator agreement. To the best of our knowledge, this is the first subjectivity dataset for the Czech language. We also created an additional dataset that consists of 200k automatically labeled sentences. Both datasets are freely available for research purposes. Furthermore, we fine-tune five pre-trained BERT-like models to set a monolingual baseline for the new dataset and we achieve 93.56% of accuracy. We fine-tune models on the existing English dataset for which we obtained results that are on par with the current state-of-the-art results. Finally, we perform zero-shot cross-lingual subjectivity classification between Czech and English to verify the usability of our dataset as the cross-lingual benchmark. We compare and discuss the cross-lingual and monolingual results and the ability of multilingual models to transfer knowledge between languages.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

pauli31/czech-subjectivity-dataset officialmentioned in papermentioned on GitHub report
pauli31/linear-transformation-4-cs-sa mentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

ClassificationSubjectivity Analysis

Datasets

Introduced by this paper, per the archive.

Czech Subjectivity Dataset

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Subjectivity Analysis Czech Subjectivity Dataset XLM-R-Large Accuracy 93.56 #1 of 5 Archive leaderboard report
Subjectivity Analysis Czech Subjectivity Dataset RobeCzech Accuracy 93.29 #2 of 5 Archive leaderboard report
Subjectivity Analysis Czech Subjectivity Dataset Czert-B Accuracy 92.85 #3 of 5 Archive leaderboard report
Subjectivity Analysis Czech Subjectivity Dataset Czech Electra Accuracy 91.85 #4 of 5 Archive leaderboard report
Subjectivity Analysis Czech Subjectivity Dataset mBERT Accuracy 91.23 #5 of 5 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections