{"url":"/dataset/smc-text-corpus","name":"SMC Text Corpus","full_name":null,"description_markdown":"Contents (As on March 4, 2019)\r\n--------\r\nThe text corpus contains running text from various free licensed sources.\r\n- The whole content of Malayalam Wikipedia extracted on January 1, 2019\r\n- News/Article from various sources, source mentioned in respective files: \r\n- 251 Mb\r\n- 8,60,159 lines\r\n- 98,15,533 words\r\n- 10,11,11,885 characters\r\n\r\nThe word corpus contains\r\n- Classified lexicon prepared for [Malaylam Morphology Analyser project](https://gitlab.com/smc/mlmorph)\r\n- Unique words extracted from Malayalam Wikipedia, Wictionary etc.\r\n- 14,27,392 words","description_withheld":null,"homepage":"https://gitlab.com/smc/corpus/-/tree/master/","introduced_date":"2019-03-02","introduced_date_note":null,"introduced_by":null,"license":{"name":"Creative Commons Attribution-ShareAlike","url":"https://creativecommons.org/licenses/by-sa/3.0/"},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Language Modelling","url":"/task/language-modelling","datasets_with_task":"/datasets/task/language-modelling"}],"languages":[{"name":"Malayalam","url":"/datasets/language/malayalam"}],"variants":["SMC Text Corpus"],"data_loaders":[],"num_papers_in_archive":2,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}