Datasets › Govdocs1

Govdocs1

17 Aug 2009 archive 2025-07-28

GovDocs is a corpus of nearly 1 million documents that are freely available for research and may be, to the best of the authors' knowledge, freely redistributed. These documents were obtained by performing searches for words randomly chosen from the Unix dictionary, numbers randomly chosen between 1 and 1 million, and randomized combinations of the two, for documents of specified file types that resided on web servers in the .gov domain using the Yahoo an Google search engines. The documents are representative of a diverse sample of real-world files of various formats produced by a variety of tools, including any malware that may be present in the files. Therefore, the corpus has been used in digital forensics, malware analysis, computer vision, and natural language processing research.

Benchmarks archive 2025-07-28

No leaderboard in the archive resolves to this dataset.

Papers archive 2025-07-28

No paper in the archive has a leaderboard row on this dataset; the archive counts 1 paper for it but never published that list.

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

No task tagged in the archive.

License archive 2025-07-28

Public Domain

Modalities archive 2025-07-28

No modality tagged.

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • Govdocs1

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections