{"url":"/dataset/indiccorp","name":"IndicCorp","full_name":null,"description_markdown":"IndicCorp is a large monolingual corpora with around 9 billion tokens covering 12 of the major Indian languages. It has been developed by discovering and scraping thousands of web sources - primarily news, magazines and books, over a duration of several months.\r\n\r\n**Languages covered**: Assamese, Bengali, English, Gujarati, Hindi, Kannada, Malayalam, Marathi, Oriya, Punjabi, Tamil, Telugu\r\n\r\n**Corpus Format**: The corpus is a single large text file containing one sentence per line. The publicly released version is randomly shuffled, untokenized and deduplicated. \r\n\r\n**Downloads**\r\n\r\n| Language | \\# News Articles* | Sentences     | Tokens        | Link     |\r\n| -------- | ----------------- | ------------- | ------------- | -------- |\r\n| as       | 0.60M             | 1.39M   |  32.6M  | [link](https://storage.googleapis.com/ai4bharat-public-indic-nlp-corpora/indiccorp/as.tar.xz) |\r\n| bn       | 3.83M             | 39.9M | 836M  | [link](https://storage.googleapis.com/ai4bharat-public-indic-nlp-corpora/indiccorp/bn.tar.xz) |\r\n| en       | 3.49M             | 54.3M | 1.22B | [link](https://storage.googleapis.com/ai4bharat-public-indic-nlp-corpora/indiccorp/en.tar.xz) |\r\n| gu       | 2.63M             | 41.1M | 719M  | [link](https://storage.googleapis.com/ai4bharat-public-indic-nlp-corpora/indiccorp/gu.tar.xz) |\r\n| hi       | 4.95M             | 63.1M |  1.86B | [link](https://storage.googleapis.com/ai4bharat-public-indic-nlp-corpora/indiccorp/hi.tar.xz) |\r\n| kn       | 3.76M             | 53.3M | 713M  | [link](https://storage.googleapis.com/ai4bharat-public-indic-nlp-corpora/indiccorp/bn.tar.xz) |\r\n| ml       | 4.75M             | 50.2M |  721M  | [link](https://storage.googleapis.com/ai4bharat-public-indic-nlp-corpora/indiccorp/ml.tar.xz) |\r\n| mr       | 2.31M             | 34.0M | 551M  | [link](https://storage.googleapis.com/ai4bharat-public-indic-nlp-corpora/indiccorp/mr.tar.xz) |\r\n| or       | 0.69M             | 6.94M   | 107M   | [link](https://storage.googleapis.com/ai4bharat-public-indic-nlp-corpora/indiccorp/or.tar.xz) |\r\n| pa       | 2.64M             | 29.2M |  773M  | [link](https://storage.googleapis.com/ai4bharat-public-indic-nlp-corpora/indiccorp/pa.tar.xz) |\r\n| ta       | 4.41M             |  31.5M   |  582M  | [link](https://storage.googleapis.com/ai4bharat-public-indic-nlp-corpora/indiccorp/ta.tar.xz) |\r\n| te       | 3.98M             | 47.9M   |  674M  | [link](https://storage.googleapis.com/ai4bharat-public-indic-nlp-corpora/indiccorp/te.tar.xz) |\r\n\r\n\\* Excluding articles obtained from the OSCAR corpus","description_withheld":null,"homepage":"https://indicnlp.ai4bharat.org/corpora/","introduced_date":"2020-11-08","introduced_date_note":null,"introduced_by":{"paper":"/paper/indicnlpsuite-monolingual-corpora-evaluation","title":"IndicNLPSuite: Monolingual Corpora, Evaluation Benchmarks and Pre-trained Multilingual Language Models for Indian Languages","first_author":"Divyanshu Kakwani","url":null},"license":null,"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[],"languages":[{"name":"English","url":"/datasets/language/english"},{"name":"Bengali","url":"/datasets/language/bengali"},{"name":"Hindi","url":"/datasets/language/hindi"},{"name":"Marathi","url":"/datasets/language/marathi"},{"name":"Tamil","url":"/datasets/language/tamil"},{"name":"Telugu","url":"/datasets/language/telugu"},{"name":"Assamese","url":"/datasets/language/assamese"},{"name":"Gujarati","url":"/datasets/language/gujarati"},{"name":"Kannada","url":"/datasets/language/kannada"},{"name":"Malayalam","url":"/datasets/language/malayalam"},{"name":"Oriya (macrolanguage)","url":"/datasets/language/oriya-macrolanguage"},{"name":"Punjabi","url":"/datasets/language/punjabi"}],"variants":["IndicCorp"],"data_loaders":[],"num_papers_in_archive":28,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}