Papers › Climate Research Domain BERTs: Pretraining, Adaptation, and Evaluation

Climate Research Domain BERTs: Pretraining, Adaptation, and Evaluation

19 May 2025Preprint 2025 5archive 2025-07-28

Andrija Poleksić, Sanda Martinčić-Ipšić

Motivated by the pressing issue of climate change and the growing volume of data, we pretrain three new language models using climate change research papers published in top-tier journals. Adaptation of existing domain-specific models is utilized for CliSciBERT and SciClimateBERT and pretraining from scratch resulted in CliReBERT (Climate Research BERT). The performance assessment is performed on the climate change NLP benchmark ClimaBench. We evaluate SciBERT, ClimateBERT, BERT, RoBERTa and DistilRoBERTa - along with our new models - CliReBERT, CliSciBERT and SciClimateBERT - using five different random seeds on all seven ClimaBench datasets. CliReBERT achieves the highest overall performance with a macro-averaged F1 score of 65.45%, and outperforms all other models on three of the seven tasks. Additionally, CliReBERT demonstrates the most stable fine-tuning behavior, yielding the lowest average standard deviation across seeds (0.0118). The 5-fold stratified cross-validation on the SciDCC dataset showed that CliReBERT achieved the highest overall macro-average F1 score (53.75%), slightly outperforming RoBERTa and DistilRoBERTa, while the domain-adapted models underperformed their base counterparts. The superior performance of CliReBERT is accompanied by the lowest tokenizer fertility, suggesting appropriateness to model domain-specific vocabulary.

PaperPDF

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Text Classification

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Text Classification Climabench CliReBERT (P0L3/clirebert_clirevocab_uncased) Evaluation Macro F1 0.6545 #1 of 8 Archive leaderboard report
Text Classification Climabench ClimateBERT (climatebert/distilroberta-base-climate-f) Evaluation Macro F1 0.6437 #2 of 8 Archive leaderboard report
Text Classification Climabench BERT (google-bert/bert-base-uncased) Evaluation Macro F1 0.6139 #3 of 8 Archive leaderboard report
Text Classification Climabench CliSciBERT (P0L3/cliscibert_scivocab_uncased) Evaluation Macro F1 0.6050 #4 of 8 Archive leaderboard report
Text Classification Climabench SciBERT (allenai/scibert_scivocab_cased) Evaluation Macro F1 0.5931 #5 of 8 Archive leaderboard report
Text Classification Climabench DistilRoBERTa (distilbert/distilroberta-base) Evaluation Macro F1 0.5835 #6 of 8 Archive leaderboard report
Text Classification Climabench SciClimateBERT (P0L3/sciclimatebert) Evaluation Macro F1 0.5783 #7 of 8 Archive leaderboard report
Text Classification Climabench RoBERTa (FacebookAI/roberta-base) Evaluation Macro F1 0.5739 #8 of 8 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

AdamAttentionAttention DropoutBASEBERTDense ConnectionsDropoutLayer NormalizationLinear LayerLinear Warmup With Linear DecayMulti-Head AttentionResidual ConnectionRoBERTaSoftmaxWeight DecayWordPiece

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections