{"url":"/dataset/moroco","name":"MOROCO","full_name":"MOldavian and ROmanian Dialectal COrpus","description_markdown":"The MOldavian and ROmanian Dialectal COrpus (MOROCO) is a corpus that contains 33,564 samples of text (with over 10 million tokens) collected from the news domain. The samples belong to one of the following six topics: culture, finance, politics, science, sports and tech. The data set is divided into 21,719 samples for training, 5,921 samples for validation and another 5,924 samples for testing. \r\n\r\nSource: [MOROCO: The Moldavian and Romanian Dialectal Corpus](/paper/moroco-the-moldavian-and-romanian-dialectal)","description_withheld":null,"homepage":"https://github.com/butnaruandrei/MOROCO","introduced_date":null,"introduced_date_note":null,"introduced_by":{"paper":"/paper/moroco-the-moldavian-and-romanian-dialectal","title":"MOROCO: The Moldavian and Romanian Dialectal Corpus","first_author":"Andrei M. Butnaru","url":null},"license":null,"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Language Identification","url":"/task/language-identification","datasets_with_task":"/datasets/task/language-identification"}],"languages":[{"name":"Romanian","url":"/datasets/language/romanian"}],"variants":["MOROCO"],"data_loaders":[{"repo":"https://github.com/huggingface/datasets","url":"https://huggingface.co/datasets/universityofbucharest/moroco","frameworks":["tf","pytorch","jax"]},{"repo":"https://github.com/huggingface/datasets","url":"https://huggingface.co/datasets/moroco","frameworks":["tf","pytorch","jax"]},{"repo":"https://github.com/butnaruandrei/MOROCO","url":"https://github.com/butnaruandrei/MOROCO","frameworks":[]}],"num_papers_in_archive":22,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}