{"url":"/dataset/shoptc-100k","name":"ShopTC-100K","full_name":null,"description_markdown":"# ShopTC-100K Dataset\r\n\r\nThe ShopTC-100K dataset is collected using [TermMiner](https://github.com/eltsai/term_miner/), an open-source data collection and topic modeling pipeline introduced in the paper: \r\n\r\n[Harmful Terms and Where to Find Them: Measuring and Modeling Unfavorable Financial Terms and Conditions in Shopping Websites at Scale](https://www.arxiv.org/abs/2502.01798)\r\n\r\nIf you find this dataset or the related paper useful for your research, please cite our paper:\r\n\r\n```\r\n@inproceedings{tsai2025harmful,\r\n  author = {Elisa Tsai and Neal Mangaokar and Boyuan Zheng and Haizhong Zheng and Atul Prakash},\r\n  title = {Harmful Terms and Where to Find Them: Measuring and Modeling Unfavorable Financial Terms and Conditions in Shopping Websites at Scale},\r\n  booktitle = {Proceedings of the ACM Web Conference 2025 (WWW ’25)},\r\n  year = {2025},\r\n  location = {Sydney, NSW, Australia},\r\n  publisher = {ACM},\r\n  address = {New York, NY, USA},\r\n  pages = {14},\r\n  month = {April 28-May 2},\r\n  doi = {10.1145/3696410.3714573}\r\n}\r\n```\r\n\r\n## Dataset Description\r\n\r\nThe dataset consists of sanitized terms extracted from e-commerce websites with English terms and conditions. The websites were sourced from the [Tranco list](https://tranco-list.eu/) (as of April 2024).","description_withheld":null,"homepage":"https://huggingface.co/datasets/eltsai/ShopTC-100K","introduced_date":"2025-02-03","introduced_date_note":null,"introduced_by":{"paper":"/paper/harmful-terms-and-where-to-find-them","title":"Harmful Terms and Where to Find Them: Measuring and Modeling Unfavorable Financial Terms and Conditions in Shopping Websites at Scale","first_author":null,"url":null},"license":{"name":"MIT","url":null},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Text Classification","url":"/task/text-classification","datasets_with_task":"/datasets/task/text-classification"},{"name":"Text Summarization","url":"/task/text-summarization","datasets_with_task":"/datasets/task/text-summarization"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["ShopTC-100K"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}