Papers › IndicNLPSuite: Monolingual Corpora, Evaluation Benchmarks and Pre-trained Multilingual...

IndicNLPSuite: Monolingual Corpora, Evaluation Benchmarks and Pre-trained Multilingual Language Models for Indian Languages

8 Nov 2020Findings of the Association for Computational Linguistics 2020archive 2025-07-28

Divyanshu Kakwani, Anoop Kunchukuttan, Satish Golla, Gokul N.C., Avik Bhattacharyya, Mitesh M. Khapra, Pratyush Kumar.

In this paper, we introduce NLP resources for 11 major Indian languages from two major language families. These resources include: (a) large-scale sentence-level monolingual corpora, (b) pre-trained word embeddings, (c) pre-trained language models, and (d) multiple NLU evaluation datasets (IndicGLUE benchmark). The monolingual corpora contains a total of 8.8 billion tokens across all 11 languages and Indian English, primarily sourced from news crawls. The word embeddings are based on FastText, hence suitable for handling morphological complexity of Indian languages. The pre-trained language models are based on the compact ALBERT model. Lastly, we compile the IndicGLUE benchmark for Indian language NLU. To this end, we create datasets for the following tasks: Article Genre Classification, Headline Prediction, Wikipedia Section-Title Prediction, Cloze-style Multiple choice QA, Winograd NLI and COPA. We also include publicly available datasets for some Indic languages for tasks like Named Entity Recognition, Cross-lingual Sentence Retrieval, Paraphrase detection, etc. Our embeddings are competitive or better than existing pre-trained embeddings on multiple tasks. We hope that the availability of the dataset will accelerate Indic NLP research which has the potential to impact more than a billion people. It can also help the community in evaluating advances in NLP over a more diverse pool of languages. The data and models are available at https://indicnlp.ai4bharat.org.

PaperPDFCode

Code

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Genre classificationMultiple Choice Question Answering (MCQA)Multiple-choiceNamed Entity RecognitionNamed Entity Recognition (NER)News ClassificationRetrievalSentenceSentence RetrievalSentiment AnalysisWord Embeddingsnamed-entity-recognition

Datasets

Introduced by this paper, per the archive.

IndicCorpIndicGLUE

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Multiple Choice Question Answering (MCQA) IndicGLUE WSTP Pa IndicBERT Large Accuracy 77.54 #2 of 3 Archive leaderboard report
News Classification Soham News Article Classification IndicBERT Base Accuracy 78.45 #3 of 3 Archive leaderboard report
Sentiment Analysis IITP Movie Reviews Sentiment IndicBERT Base Accuracy 59.03 #3 of 3 Archive leaderboard report
Sentiment Analysis IITP Product Reviews Sentiment IndicBERT Base Accuracy 71.32 #4 of 4 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

ALBERTAdamAttentionDense ConnectionsLAMBLayer NormalizationLinear LayerMulti-Head AttentionResidual ConnectionSoftmaxWordPiece

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections