{"url":"/dataset/enterprise-driven-open-source-software","name":"Enterprise-Driven Open Source Software","full_name":null,"description_markdown":"This is a dataset of open source software developed mainly by enterprises rather than volunteers. This can be used to address known generalizability concerns, and, also, to perform research on open source business software development. Based on the premise that an enterprise's employees are likely to contribute to a project developed by their organization using the email account provided by it, we mine domain names associated with enterprises from open data sources as well as through white- and blacklisting, and use them through three heuristics to identify 17,264 enterprise GitHub projects. We provide these as a dataset detailing their provenance and properties. A manual evaluation of a dataset sample shows an identification accuracy of 89%.","description_withheld":null,"homepage":"https://zenodo.org/record/3742973","introduced_date":"2020-04-21","introduced_date_note":null,"introduced_by":{"paper":null,"title":"A Dataset of Enterprise-Driven Open Source Software","first_author":null,"url":null},"license":{"name":"Apache License 2.0","url":"https://opensource.org/licenses/Apache-2.0"},"modalities":[],"tasks":[],"languages":[],"variants":["Enterprise-Driven Open Source Software"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}