{"url":"/dataset/counter","name":"COUNTER","full_name":null,"description_markdown":"The **COUNTER** (COrpus of Urdu News TExt Reuse) corpus contains 600 source-derived document pairs collected from the field of journalism. It can be used to evaluate mono-lingual text reuse detection systems in general and specifically for Urdu language.\r\n\r\nThe corpus has 600 source and 600 derived documents. It contains in total 275,387 words (tokens), 21,426 unique words and 10,841 sentences. It has been manually annotated at document level with three levels of reuse: wholly derived (135), partially derived (288) and non derived (177).\r\n\r\nSource: [COUNTER](http://ucrel.lancs.ac.uk/textreuse/counter.php)","description_withheld":null,"homepage":"http://ucrel.lancs.ac.uk/textreuse/counter.php","introduced_date":null,"introduced_date_note":null,"introduced_by":{"paper":null,"title":"COUNTER - corpus of Urdu news text reuse","first_author":null,"url":"http://dx.doi.org/10.1007/s10579-016-9367-2"},"license":null,"modalities":[],"tasks":[],"languages":[{"name":"Urdu","url":"/datasets/language/urdu"}],"variants":["COUNTER"],"data_loaders":[{"repo":"https://github.com/huggingface/datasets","url":"https://huggingface.co/datasets/ucrelnlp/counter","frameworks":["tf","pytorch","jax"]},{"repo":"https://github.com/huggingface/datasets","url":"https://huggingface.co/datasets/counter","frameworks":["tf","pytorch","jax"]}],"num_papers_in_archive":3,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}