{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/evaluating-approaches-for-supervised-semantic","title":"Evaluating approaches for supervised semantic labeling","arxiv_id":"1801.09788","date":"2018-01-29","proceeding":null,"authors":["Natalia Ruemmele","Yuriy Tyshetskiy","Alex Collins"],"abstract":"Relational data sources are still one of the most popular ways to store\nenterprise or Web data, however, the issue with relational schema is the lack\nof a well-defined semantic description. A common ontology provides a way to\nrepresent the meaning of a relational schema and can facilitate the integration\nof heterogeneous data sources within a domain. Semantic labeling is achieved by\nmapping attributes from the data sources to the classes and properties in the\nontology. We formulate this problem as a multi-class classification problem\nwhere previously labeled data sources are used to learn rules for labeling new\ndata sources. The majority of existing approaches for semantic labeling have\nfocused on data integration challenges such as naming conflicts and semantic\nheterogeneity. In addition, machine learning approaches typically have issues\naround class imbalance, lack of labeled instances and relative importance of\nattributes. To address these issues, we develop a new machine learning model\nwith engineered features as well as two deep learning models which do not\nrequire extensive feature engineering. We evaluate our new approaches with the\nstate-of-the-art.","url_abs":"http://arxiv.org/abs/1801.09788v1","url_pdf":"http://arxiv.org/pdf/1801.09788v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"evaluating-approaches-for-supervised-semantic","repo_url":"https://github.com/NICTA/serene-benchmark","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"machine-learning","task_name":"BIG-bench Machine Learning"},{"task_slug":"data-integration","task_name":"Data Integration"},{"task_slug":"feature-engineering","task_name":"Feature Engineering"},{"task_slug":"multi-class-classification","task_name":"Multi-class Classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}