{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/did-you-really-just-have-a-heart-attack","title":"Did You Really Just Have a Heart Attack? Towards Robust Detection of Personal Health Mentions in Social Media","arxiv_id":"1802.09130","date":"2018-02-26","proceeding":null,"authors":["Payam Karisani","Eugene Agichtein"],"abstract":"Millions of users share their experiences on social media sites, such as\nTwitter, which in turn generate valuable data for public health monitoring,\ndigital epidemiology, and other analyses of population health at global scale.\nThe first, critical, task for these applications is classifying whether a\npersonal health event was mentioned, which we call the (PHM) problem. This task\nis challenging for many reasons, including typically short length of social\nmedia posts, inventive spelling and lexicons, and figurative language,\nincluding hyperbole using diseases like \"heart attack\" or \"cancer\" for\nemphasis, and not as a health self-report. This problem is even more\nchallenging for rarely reported, or frequent but ambiguously expressed\nconditions, such as \"stroke\". To address this problem, we propose a general,\nrobust method for detecting PHMs in social media, which we call WESPAD, that\ncombines lexical, syntactic, word embedding-based, and context-based features.\nWESPAD is able to generalize from few examples by automatically distorting the\nword embedding space to most effectively detect the true health mentions.\nUnlike previously proposed state-of-the-art supervised and deep-learning\ntechniques, WESPAD requires relatively little training data, which makes it\npossible to adapt, with minimal effort, to each new disease and condition. We\nevaluate WESPAD on both an established publicly available Flu detection\nbenchmark, and on a new dataset that we have constructed with mentions of\nmultiple health conditions. Our experiments show that WESPAD outperforms the\nbaselines and state-of-the-art methods, especially in cases when the number and\nproportion of true health mentions in the training data is small.","url_abs":"http://arxiv.org/abs/1802.09130v2","url_pdf":"http://arxiv.org/pdf/1802.09130v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"did-you-really-just-have-a-heart-attack","repo_url":"https://github.com/emory-irlab/PHM2017","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null},{"paper_slug":"did-you-really-just-have-a-heart-attack","repo_url":"https://github.com/p-karisani/FirstPHM","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"epidemiology","task_name":"Epidemiology"},{"task_slug":"semi-supervised-text-classification-1","task_name":"Semi-Supervised Text Classification"},{"task_slug":"text-classification","task_name":"Text Classification"}],"methods":[],"datasets_introduced":[{"slug":"phm2017","name":"PHM2017","full_name":"PHM2017"}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}