Methods › Natural Language Processing › Language Models › mBERT
mBERT
Introduced by Jacob Devlin et al. in BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
mBERT
Papers archive 2025-07-28
30 shown of 198, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
How do datasets, developers, and models affect biases in a low-resourced language? 7 Jun 2025 · 0 repositories · arXiv:2506.06816
-
University of Indonesia at SemEval-2025 Task 11: Evaluating State-of-the-Art Encoders for Multi-Label Emotion Detection 22 May 2025 · 0 repositories · arXiv:2505.16460
-
Cross-Linguistic Transfer in Multilingual NLP: The Role of Language Families and Morphology 20 May 2025 · 0 repositories · arXiv:2505.13908
-
Generating Synthetic Oracle Datasets to Analyze Noise Impact: A Study on Building Function Classification Using Tweets 28 Mar 2025 · 0 repositories · arXiv:2503.22856
-
Comparative Approaches to Sentiment Analysis Using Datasets in Major European and Arabic Languages 21 Jan 2025 · 0 repositories · arXiv:2501.12540
-
The First Multilingual Model For The Detection of Suicide Texts 20 Dec 2024 · 0 repositories · arXiv:2412.15498
-
Transformer-Based Contextualized Language Models Joint with Neural Networks for Natural Language Inference in Vietnamese 20 Nov 2024 · 0 repositories · arXiv:2411.13407
-
From N-grams to Pre-trained Multilingual Models For Language Identification 11 Oct 2024 · 2 repositories · arXiv:2410.08728
-
Motamot: A Dataset for Revealing the Supremacy of Large Language Models over Transformer Models in Bengali Political Sentiment Analysis 28 Jul 2024 · 1 repository · arXiv:2407.19528
-
On Initializing Transformers with Pre-trained Embeddings 17 Jul 2024 · 0 repositories · arXiv:2407.12514
-
Decipherment-Aware Multilingual Learning in Jointly Trained Language Models 11 Jun 2024 · 0 repositories · arXiv:2406.07231
-
Vietnamese AI Generated Text Detection 6 May 2024 · 0 repositories · arXiv:2405.03206
-
Unraveling the Dominance of Large Language Models Over Transformer Models for Bangla Natural Language Inference: A Comprehensive Study 5 May 2024 · 1 repository · arXiv:2405.02937
-
UQA: Corpus for Urdu Question Answering 2 May 2024 · 3 repositories · arXiv:2405.01458
-
Incorporating Lexical and Syntactic Knowledge for Unsupervised Cross-Lingual Transfer 25 Apr 2024 · 1 repository · arXiv:2404.16627
-
Adapting Mental Health Prediction Tasks for Cross-lingual Learning via Meta-Training and In-context Learning with Large Language Model 13 Apr 2024 · 0 repositories · arXiv:2404.09045
-
Data Bias According to Bipol: Men are Naturally Right and It is the Role of Women to Follow Their Lead 7 Apr 2024 · 1 repository · arXiv:2404.04838
-
Large Language Models for Expansion of Spoken Language Understanding Systems to New Languages 3 Apr 2024 · 1 repository · arXiv:2404.02588
-
A Benchmark Evaluation of Clinical Named Entity Recognition in French 28 Mar 2024 · 0 repositories · arXiv:2403.19726
-
Introducing Syllable Tokenization for Low-resource Languages: A Case Study with Swahili 26 Mar 2024 · 0 repositories · arXiv:2406.15358
-
Ax-to-Grind Urdu: Benchmark Dataset for Urdu Fake News Detection 20 Mar 2024 · 1 repository · arXiv:2403.14037
-
LLMLingua-2: Data Distillation for Efficient and Faithful Task-Agnostic Prompt Compression 19 Mar 2024 · 1 repository · arXiv:2403.12968
-
Evaluating Named Entity Recognition: A comparative analysis of mono- and multilingual transformer models on a novel Brazilian corporate earnings call transcripts dataset 18 Mar 2024 · 2 repositories · arXiv:2403.12212
-
A Measure for Transparent Comparison of Linguistic Diversity in Multilingual NLP Data Sets 6 Mar 2024 · 1 repository · arXiv:2403.03909
-
Enhancing ESG Impact Type Identification through Early Fusion and Multilingual Models 16 Feb 2024 · 0 repositories · arXiv:2402.10772
-
Cross-lingual Editing in Multilingual Language Models 19 Jan 2024 · 1 repository · arXiv:2401.10521
-
LinguAlchemy: Fusing Typological and Geographical Elements for Unseen Language Generalization 11 Jan 2024 · 0 repositories · arXiv:2401.06034
-
Multilingual large language models leak human stereotypes across language boundaries 12 Dec 2023 · 1 repository · arXiv:2312.07141
-
On Significance of Subword tokenization for Low Resource and Efficient Named Entity Recognition: A case study in Marathi 3 Dec 2023 · 0 repositories · arXiv:2312.01306
-
Vashantor: A Large-scale Multilingual Benchmark Dataset for Automated Translation of Bangla Regional Dialects to Bangla Language 18 Nov 2023 · 1 repository · arXiv:2311.11142
Tasks archive 2025-07-28
20 shown of 170 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections