{"url":"/dataset/dataset-of-dockerfiles","name":"Dataset of Dockerfiles","full_name":null,"description_markdown":"This dataset of approximately 178,000 unique Dockerfiles collected from GitHub to facilitate sophisticated semantics-aware static analysis of Dockerfiles. To enhance the usability of this data, the authors use five representations for working with, mining from, and analyzing these Dockerfiles. Each Dockerfile representation builds upon the previous ones, and the final representation, created by three levels of nested parsing and abstraction, makes tasks such as mining and static checking tractable.","description_withheld":null,"homepage":"https://doi.org/10.5281/zenodo.3628771","introduced_date":"2020-03-28","introduced_date_note":null,"introduced_by":{"paper":null,"title":"A Dataset of Dockerfiles","first_author":null,"url":null},"license":{"name":"MIT License","url":"https://opensource.org/licenses/MIT"},"modalities":[],"tasks":[],"languages":[],"variants":["Dataset of Dockerfiles"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}