{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-practitioners-guide-to-transfer-learning","title":"A Practitioners' Guide to Transfer Learning for Text Classification using Convolutional Neural Networks","arxiv_id":"1801.06480","date":"2018-01-19","proceeding":null,"authors":["Tushar Semwal","Gaurav Mathur","Promod Yenigalla","Shivashankar B. Nair"],"abstract":"Transfer Learning (TL) plays a crucial role when a given dataset has\ninsufficient labeled examples to train an accurate model. In such scenarios,\nthe knowledge accumulated within a model pre-trained on a source dataset can be\ntransferred to a target dataset, resulting in the improvement of the target\nmodel. Though TL is found to be successful in the realm of image-based\napplications, its impact and practical use in Natural Language Processing (NLP)\napplications is still a subject of research. Due to their hierarchical\narchitecture, Deep Neural Networks (DNN) provide flexibility and customization\nin adjusting their parameters and depth of layers, thereby forming an apt area\nfor exploiting the use of TL. In this paper, we report the results and\nconclusions obtained from extensive empirical experiments using a Convolutional\nNeural Network (CNN) and try to uncover thumb rules to ensure a successful\npositive transfer. In addition, we also highlight the flawed means that could\nlead to a negative transfer. We explore the transferability of various layers\nand describe the effect of varying hyper-parameters on the transfer\nperformance. Also, we present a comparison of accuracy value and model size\nagainst state-of-the-art methods. Finally, we derive inferences from the\nempirical results and provide best practices to achieve a successful positive\ntransfer.","url_abs":"http://arxiv.org/abs/1801.06480v1","url_pdf":"http://arxiv.org/pdf/1801.06480v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-practitioners-guide-to-transfer-learning","repo_url":"https://github.com/tushar-semwal/TransferLearning_CNN_TextClassification","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"text-classification","task_name":"Text Classification"},{"task_slug":"transfer-learning","task_name":"Transfer Learning"},{"task_slug":"text-classification-1","task_name":"text-classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}