{"url":"/dataset/various-url-datasets","name":"Various URL Datasets","full_name":"https://github.com/ada-url/url-various-datasets","description_markdown":"## Various URL Datasets\r\n\r\nThese are collections of URLs for  benchmarking purposes.\r\n\r\n- files/node_files.txt: all source files from a given Node.js snapshot as URLs (43415 URLs).\r\n- files/linux_files.txt: all files from a Linux systems as URLs (169312 URLs).\r\n- wikipedia/wikipedia_100k.txt: 100k URLs from a snapshot of all Wikipedia articles as URLs (March 6th 2023)\r\n- others/kasztp.txt: test URLs from https://github.com/kasztp/URL_Shortener (MIT License) (48009 URLs).\r\n- others/userbait.txt : test URLs from https://github.com/userbait/phishing_sites_detector (unknown copyright) (11430 URLs).\r\n- top100/top100.txt: crawl of  the top visited 100 websites and extracts unique URLs\r\n\r\n**Disclaimer**: This repository is developed and released for research purposes only. \r\n  - This project reshares some publicly available datasets. When in doubt, investigate the copyright of the files you want to use. \r\n  - There may be errors and duplicates in these files.","description_withheld":null,"homepage":"","introduced_date":"2023-11-17","introduced_date_note":null,"introduced_by":{"paper":"/paper/parsing-millions-of-urls-per-second","title":"Parsing Millions of URLs per Second","first_author":null,"url":null},"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["Various URL Datasets"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}