{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/snorkel-rapid-training-data-creation-with","title":"Snorkel: Rapid Training Data Creation with Weak Supervision","arxiv_id":"1711.10160","date":"2017-11-28","proceeding":null,"authors":["Alexander Ratner","Stephen H. Bach","Henry Ehrenberg","Jason Fries","Sen Wu","Christopher Ré"],"abstract":"Labeling training data is increasingly the largest bottleneck in deploying\nmachine learning systems. We present Snorkel, a first-of-its-kind system that\nenables users to train state-of-the-art models without hand labeling any\ntraining data. Instead, users write labeling functions that express arbitrary\nheuristics, which can have unknown accuracies and correlations. Snorkel\ndenoises their outputs without access to ground truth by incorporating the\nfirst end-to-end implementation of our recently proposed machine learning\nparadigm, data programming. We present a flexible interface layer for writing\nlabeling functions based on our experience over the past year collaborating\nwith companies, agencies, and research labs. In a user study, subject matter\nexperts build models 2.8x faster and increase predictive performance an average\n45.5% versus seven hours of hand labeling. We study the modeling tradeoffs in\nthis new setting and propose an optimizer for automating tradeoff decisions\nthat gives up to 1.8x speedup per pipeline execution. In two collaborations,\nwith the U.S. Department of Veterans Affairs and the U.S. Food and Drug\nAdministration, and on four open-source text and image data sets representative\nof other deployments, Snorkel provides 132% average improvements to predictive\nperformance over prior heuristic approaches and comes within an average 3.60%\nof the predictive performance of large hand-curated training sets.","url_abs":"http://arxiv.org/abs/1711.10160v1","url_pdf":"http://arxiv.org/pdf/1711.10160v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"snorkel-rapid-training-data-creation-with","repo_url":"https://github.com/HazyResearch/metal","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"snorkel-rapid-training-data-creation-with","repo_url":"https://github.com/megagonlabs/ruler","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"machine-learning","task_name":"BIG-bench Machine Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1711.10160","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}