Datasets › Various URL Datasets

Various URL Datasets (https://github.com/ada-url/url-various-datasets)

Introduced in Parsing Millions of URLs per Second17 Nov 2023 archive 2025-07-28

Various URL Datasets

These are collections of URLs for benchmarking purposes.

  • files/node_files.txt: all source files from a given Node.js snapshot as URLs (43415 URLs).
  • files/linux_files.txt: all files from a Linux systems as URLs (169312 URLs).
  • wikipedia/wikipedia_100k.txt: 100k URLs from a snapshot of all Wikipedia articles as URLs (March 6th 2023)
  • others/kasztp.txt: test URLs from https://github.com/kasztp/URL_Shortener (MIT License) (48009 URLs).
  • others/userbait.txt : test URLs from https://github.com/userbait/phishing_sites_detector (unknown copyright) (11430 URLs).
  • top100/top100.txt: crawl of the top visited 100 websites and extracts unique URLs

Disclaimer: This repository is developed and released for research purposes only. - This project reshares some publicly available datasets. When in doubt, investigate the copyright of the files you want to use. - There may be errors and duplicates in these files.

Benchmarks archive 2025-07-28

No leaderboard in the archive resolves to this dataset.

Papers archive 2025-07-28

No paper in the archive has a leaderboard row on this dataset; the archive counts 1 paper for it but never published that list.

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

No task tagged in the archive.

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

No modality tagged.

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • Various URL Datasets

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections