{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/navigating-dataset-documentations-in-ai-a","title":"Navigating Dataset Documentations in AI: A Large-Scale Analysis of Dataset Cards on Hugging Face","arxiv_id":"2401.13822","date":"2024-01-24","proceeding":null,"authors":["Xinyu Yang","Weixin Liang","James Zou"],"abstract":"Advances in machine learning are closely tied to the creation of datasets. While data documentation is widely recognized as essential to the reliability, reproducibility, and transparency of ML, we lack a systematic empirical understanding of current dataset documentation practices. To shed light on this question, here we take Hugging Face -- one of the largest platforms for sharing and collaborating on ML models and datasets -- as a prominent case study. By analyzing all 7,433 dataset documentation on Hugging Face, our investigation provides an overview of the Hugging Face dataset ecosystem and insights into dataset documentation practices, yielding 5 main findings: (1) The dataset card completion rate shows marked heterogeneity correlated with dataset popularity. (2) A granular examination of each section within the dataset card reveals that the practitioners seem to prioritize Dataset Description and Dataset Structure sections, while the Considerations for Using the Data section receives the lowest proportion of content. (3) By analyzing the subsections within each section and utilizing topic modeling to identify key topics, we uncover what is discussed in each section, and underscore significant themes encompassing both technical and social impacts, as well as limitations within the Considerations for Using the Data section. (4) Our findings also highlight the need for improved accessibility and reproducibility of datasets in the Usage sections. (5) In addition, our human annotation evaluation emphasizes the pivotal role of comprehensive dataset content in shaping individuals' perceptions of a dataset card's overall quality. Overall, our study offers a unique perspective on analyzing dataset documentation through large-scale data science analysis and underlines the need for more thorough dataset documentation in machine learning research.","url_abs":"https://arxiv.org/abs/2401.13822v1","url_pdf":"https://arxiv.org/pdf/2401.13822v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"navigating-dataset-documentations-in-ai-a","repo_url":"https://github.com/youngxinyu1802/huggingface-dataset-card-analysis","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2401.13822","atlas_url":"https://app.syntology.ai/?focus=2401.13822","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2401.13822"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/huggingface/datasets","reach":{"status":"ok","spdx":"Apache-2.0"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/youngxinyu1802/huggingface-dataset-card-analysis","reach":{"status":"ok"}}],"summary":{"ran":3,"unverified":3},"by_repo_kind":{"found_in_text":{"samples":6,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"c915bf78a27d08b3","entry":"contains_wildcards","repo":"huggingface/datasets","repo_kind":"found_in_text","path":"src/datasets/data_files.py","file_url":"https://github.com/huggingface/datasets/blob/HEAD/src/datasets/data_files.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"c915bf78a27d08b3"}},{"code_sha256_prefix":"fdde5c1fc2fe5556","entry":"generate_random_fingerprint","repo":"huggingface/datasets","repo_kind":"found_in_text","path":"src/datasets/fingerprint.py","file_url":"https://github.com/huggingface/datasets/blob/HEAD/src/datasets/fingerprint.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"fdde5c1fc2fe5556"}},{"code_sha256_prefix":"91abb27e4b2dea99","entry":"transmit_format","repo":"huggingface/datasets","repo_kind":"found_in_text","path":"src/datasets/arrow_dataset.py","file_url":"https://github.com/huggingface/datasets/blob/HEAD/src/datasets/arrow_dataset.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"91abb27e4b2dea99"}},{"code_sha256_prefix":"2074bb03247a0985","entry":"delete_from_hub","repo":"huggingface/datasets","repo_kind":"found_in_text","path":"src/datasets/hub.py","file_url":"https://github.com/huggingface/datasets/blob/HEAD/src/datasets/hub.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"2074bb03247a0985"}},{"code_sha256_prefix":"c7e8b6cb6ebf0f3d","entry":"generate_fingerprint","repo":"huggingface/datasets","repo_kind":"found_in_text","path":"src/datasets/fingerprint.py","file_url":"https://github.com/huggingface/datasets/blob/HEAD/src/datasets/fingerprint.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"c7e8b6cb6ebf0f3d"}},{"code_sha256_prefix":"8e3fe703e41ff602","entry":"get_writer_batch_size_from_data_size","repo":"huggingface/datasets","repo_kind":"found_in_text","path":"src/datasets/arrow_writer.py","file_url":"https://github.com/huggingface/datasets/blob/HEAD/src/datasets/arrow_writer.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"8e3fe703e41ff602"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}