{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/ai4bharat-indicnlp-corpus-monolingual-corpora","title":"AI4Bharat-IndicNLP Corpus: Monolingual Corpora and Word Embeddings for Indic Languages","arxiv_id":"2005.00085","date":"2020-04-30","proceeding":null,"authors":["Anoop Kunchukuttan","Divyanshu Kakwani","Satish Golla","Gokul N. C.","Avik Bhattacharyya","Mitesh M. Khapra","Pratyush Kumar"],"abstract":"We present the IndicNLP corpus, a large-scale, general-domain corpus containing 2.7 billion words for 10 Indian languages from two language families. We share pre-trained word embeddings trained on these corpora. We create news article category classification datasets for 9 languages to evaluate the embeddings. We show that the IndicNLP embeddings significantly outperform publicly available pre-trained embedding on multiple evaluation tasks. We hope that the availability of the corpus will accelerate Indic NLP research. The resources are available at https://github.com/ai4bharat-indicnlp/indicnlp_corpus.","url_abs":"https://arxiv.org/abs/2005.00085v1","url_pdf":"https://arxiv.org/pdf/2005.00085v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"ai4bharat-indicnlp-corpus-monolingual-corpora","repo_url":"https://github.com/ai4bharat-indicnlp/indicnlp_corpus","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"ai4bharat-indicnlp-corpus-monolingual-corpora","repo_url":"https://github.com/csebuetnlp/banglabert","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"word-embeddings","task_name":"Word Embeddings"}],"methods":[],"datasets_introduced":[{"slug":"indicnlp-corpus","name":"IndicNLP Corpus","full_name":null}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2005.00085","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}