{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/crowdsourcing-ground-truth-for-medical","title":"Crowdsourcing Ground Truth for Medical Relation Extraction","arxiv_id":"1701.02185","date":"2017-01-09","proceeding":null,"authors":["Anca Dumitrache","Lora Aroyo","Chris Welty"],"abstract":"Cognitive computing systems require human labeled data for evaluation, and\noften for training. The standard practice used in gathering this data minimizes\ndisagreement between annotators, and we have found this results in data that\nfails to account for the ambiguity inherent in language. We have proposed the\nCrowdTruth method for collecting ground truth through crowdsourcing, that\nreconsiders the role of people in machine learning based on the observation\nthat disagreement between annotators provides a useful signal for phenomena\nsuch as ambiguity in the text. We report on using this method to build an\nannotated data set for medical relation extraction for the $cause$ and $treat$\nrelations, and how this data performed in a supervised training experiment. We\ndemonstrate that by modeling ambiguity, labeled data gathered from crowd\nworkers can (1) reach the level of quality of domain experts for this task\nwhile reducing the cost, and (2) provide better training data at scale than\ndistant supervision. We further propose and validate new weighted measures for\nprecision, recall, and F-measure, that account for ambiguity in both human and\nmachine performance on this task.","url_abs":"http://arxiv.org/abs/1701.02185v2","url_pdf":"http://arxiv.org/pdf/1701.02185v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"crowdsourcing-ground-truth-for-medical","repo_url":"https://github.com/CrowdTruth/Medical-Relation-Extraction","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"medical-relation-extraction","task_name":"Medical Relation Extraction"},{"task_slug":null,"task_name":"Relation"},{"task_slug":"relation-extraction","task_name":"Relation Extraction"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1701.02185","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}