Papers › CalBERT - Code-mixed Adaptive Language representations using BERT

CalBERT - Code-mixed Adaptive Language representations using BERT

14 Apr 2022AAAI-MAKE 2022 4archive 2025-07-28

Aditeya Baral, Ansh Sarkar, Aronya Baksy, Deeksha D, Ashwini M Joshi

A code-mixed language is a type of language that involves the combination of two or more language varieties in its script or speech. Analysis of code-text is difficult to tackle because the language present is not consistent and does not work with existing monolingual approaches. We propose a novel approach to improve performance in Transformers by introducing an additional step called "Siamese Pre-Training", which allows pre-trained monolingual Transformers to adapt language representations for code-mixed languages with a few examples of code-mixed data. The proposed architectures beat the state of the art F1-score on the Sentiment Analysis for Indian Languages (SAIL) dataset, with the highest possible improvement being 5.1 points, while also achieving the state-of-the-art accuracy on the IndicGLUE Product Reviews dataset by beating the benchmark by 0.4 points.

PaperPDFCode

Code

aditeyabaral/calbert officialmentioned in paperpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Natural Language UnderstandingSentiment Analysis

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Sentiment Analysis IITP Product Reviews Sentiment CalBERT Accuracy 79.4 #1 of 4 Archive leaderboard report
Sentiment Analysis SAIL 2017 CalBERT F1 62 #1 of 1 Archive leaderboard report
Sentiment Analysis SAIL 2017 CalBERT Precision 61.8 #1 of 1 Archive leaderboard report
Sentiment Analysis SAIL 2017 CalBERT Recall 61.8 #1 of 1 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

AdamAttentionAttention DropoutBERTContrastive LearningDense ConnectionsDropoutLayer NormalizationLinear LayerLinear Warmup With Linear DecayMulti-Head AttentionResidual ConnectionSoftmaxWeight DecayWordPiece

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections