{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/post-training-in-deep-learning-with-last","title":"Post Training in Deep Learning with Last Kernel","arxiv_id":"1611.04499","date":"2016-11-14","proceeding":null,"authors":["Thomas Moreau","Julien Audiffren"],"abstract":"One of the main challenges of deep learning methods is the choice of an\nappropriate training strategy. In particular, additional steps, such as\nunsupervised pre-training, have been shown to greatly improve the performances\nof deep structures. In this article, we propose an extra training step, called\npost-training, which only optimizes the last layer of the network. We show that\nthis procedure can be analyzed in the context of kernel theory, with the first\nlayers computing an embedding of the data and the last layer a statistical\nmodel to solve the task based on this embedding. This step makes sure that the\nembedding, or representation, of the data is used in the best possible way for\nthe considered task. This idea is then tested on multiple architectures with\nvarious data sets, showing that it consistently provides a boost in\nperformance.","url_abs":"http://arxiv.org/abs/1611.04499v2","url_pdf":"http://arxiv.org/pdf/1611.04499v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"post-training-in-deep-learning-with-last","repo_url":"https://github.com/tomMoral/post_training","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"deep-learning","task_name":"Deep Learning"},{"task_slug":"unsupervised-pre-training","task_name":"Unsupervised Pre-training"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}