{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/ballpark-crowdsourcing-the-wisdom-of-rough","title":"Ballpark Crowdsourcing: The Wisdom of Rough Group Comparisons","arxiv_id":"1712.04828","date":"2017-12-13","proceeding":null,"authors":["Tom Hope","Dafna Shahaf"],"abstract":"Crowdsourcing has become a popular method for collecting labeled training\ndata. However, in many practical scenarios traditional labeling can be\ndifficult for crowdworkers (for example, if the data is high-dimensional or\nunintuitive, or the labels are continuous).\n  In this work, we develop a novel model for crowdsourcing that can complement\nstandard practices by exploiting people's intuitions about groups and relations\nbetween them. We employ a recent machine learning setting, called Ballpark\nLearning, that can estimate individual labels given only coarse, aggregated\nsignal over groups of data points. To address the important case of continuous\nlabels, we extend the Ballpark setting (which focused on classification) to\nregression problems. We formulate the problem as a convex optimization problem\nand propose fast, simple methods with an innate robustness to outliers.\n  We evaluate our methods on real-world datasets, demonstrating how useful\nconstraints about groups can be harnessed from a crowd of non-experts. Our\nmethods can rival supervised models trained on many true labels, and can obtain\nconsiderably better results from the crowd than a standard label-collection\nprocess (for a lower price). By collecting rough guesses on groups of instances\nand using machine learning to infer the individual labels, our lightweight\nframework is able to address core crowdsourcing challenges and train machine\nlearning models in a cost-effective way.","url_abs":"http://arxiv.org/abs/1712.04828v1","url_pdf":"http://arxiv.org/pdf/1712.04828v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"ballpark-crowdsourcing-the-wisdom-of-rough","repo_url":"https://github.com/ttthhh/ballpark","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"machine-learning","task_name":"BIG-bench Machine Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}