{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/datasets-a-community-library-for-natural","title":"Datasets: A Community Library for Natural Language Processing","arxiv_id":"2109.02846","date":"2021-09-07","proceeding":"EMNLP (ACL) 2021 11","authors":["Quentin Lhoest","Albert Villanova del Moral","Yacine Jernite","Abhishek Thakur","Patrick von Platen","Suraj Patil","Julien Chaumond","Mariama Drame","Julien Plu","Lewis Tunstall","Joe Davison","Mario Šaško","Gunjan Chhablani","Bhavitvya Malik","Simon Brandeis","Teven Le Scao","Victor Sanh","Canwen Xu","Nicolas Patry","Angelina McMillan-Major","Philipp Schmid","Sylvain Gugger","Clément Delangue","Théo Matussière","Lysandre Debut","Stas Bekman","Pierric Cistac","Thibault Goehringer","Victor Mustar","François Lagunas","Alexander M. Rush","Thomas Wolf"],"abstract":"The scale, variety, and quantity of publicly-available NLP datasets has grown rapidly as researchers propose new tasks, larger models, and novel benchmarks. Datasets is a community library for contemporary NLP designed to support this ecosystem. Datasets aims to standardize end-user interfaces, versioning, and documentation, while providing a lightweight front-end that behaves similarly for small datasets as for internet-scale corpora. The design of the library incorporates a distributed, community-driven approach to adding datasets and documenting usage. After a year of development, the library now includes more than 650 unique datasets, has more than 250 contributors, and has helped support a variety of novel cross-dataset research projects and shared tasks. The library is available at https://github.com/huggingface/datasets.","url_abs":"https://arxiv.org/abs/2109.02846v1","url_pdf":"https://arxiv.org/pdf/2109.02846v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"datasets-a-community-library-for-natural","repo_url":"https://github.com/huggingface/datasets","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"image-classification","task_name":"Image Classification"},{"task_slug":"object-recognition","task_name":"Object Recognition"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"time-series","task_name":"Time Series Analysis"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2109.02846","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}