{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/semantic-classification-of-tabular-datasets","title":"Semantic Classification of Tabular Datasets via Character-Level Convolutional Neural Networks","arxiv_id":"1901.08456","date":"2019-01-24","proceeding":null,"authors":["Paul Azunre","Craig Corcoran","Numa Dhamani","Jeffrey Gleason","Garrett Honke","David Sullivan","Rebecca Ruppel","Sandeep Verma","Jonathon Morgan"],"abstract":"A character-level convolutional neural network (CNN) motivated by\napplications in \"automated machine learning\" (AutoML) is proposed to\nsemantically classify columns in tabular data. Simulated data containing a set\nof base classes is first used to learn an initial set of weights. Hand-labeled\ndata from the CKAN repository is then used in a transfer-learning paradigm to\nadapt the initial weights to a more sophisticated representation of the problem\n(e.g., including more classes). In doing so, realistic data imperfections are\nlearned and the set of classes handled can be expanded from the base set with\nreduced labeled data and computing power requirements. Results show the\neffectiveness and flexibility of this approach in three diverse domains:\nsemantic classification of tabular data, age prediction from social media\nposts, and email spam classification. In addition to providing further evidence\nof the effectiveness of transfer learning in natural language processing (NLP),\nour experiments suggest that analyzing the semantic structure of language at\nthe character level without additional metadata---i.e., network structure,\nheaders, etc.---can produce competitive accuracy for type classification, spam\nclassification, and social media age prediction. We present our open-source\ntoolkit SIMON, an acronym for Semantic Inference for the Modeling of\nONtologies, which implements this approach in a user-friendly and\nscalable/parallelizable fashion.","url_abs":"http://arxiv.org/abs/1901.08456v1","url_pdf":"http://arxiv.org/pdf/1901.08456v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"semantic-classification-of-tabular-datasets","repo_url":"https://github.com/algorine/nokore","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"automl","task_name":"AutoML"},{"task_slug":"classification-1","task_name":"Classification"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"transfer-learning","task_name":"Transfer Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}