{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-framework-for-information-extraction-from","title":"A framework for information extraction from tables in biomedical literature","arxiv_id":"1902.10031","date":"2019-02-26","proceeding":null,"authors":["Nikola Milosevic","Cassie Gregson","Robert Hernandez","Goran Nenadic"],"abstract":"The scientific literature is growing exponentially, and professionals are no\nmore able to cope with the current amount of publications. Text mining provided\nin the past methods to retrieve and extract information from text; however,\nmost of these approaches ignored tables and figures. The research done in\nmining table data still does not have an integrated approach for mining that\nwould consider all complexities and challenges of a table. Our research is\nexamining the methods for extracting numerical (number of patients, age, gender\ndistribution) and textual (adverse reactions) information from tables in the\nclinical literature. We present a requirement analysis template and an integral\nmethodology for information extraction from tables in clinical domain that\ncontains 7 steps: (1) table detection, (2) functional processing, (3)\nstructural processing, (4) semantic tagging, (5) pragmatic processing, (6) cell\nselection and (7) syntactic processing and extraction. Our approach performed\nwith the F-measure ranged between 82 and 92%, depending on the variable, task\nand its complexity.","url_abs":"http://arxiv.org/abs/1902.10031v1","url_pdf":"http://arxiv.org/pdf/1902.10031v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-framework-for-information-extraction-from","repo_url":"https://github.com/nikolamilosevic86/TabInOut","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"table-detection","task_name":"Table Detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1902.10031","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}