{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/masters-thesis-deep-learning-for-visual","title":"Master's Thesis : Deep Learning for Visual Recognition","arxiv_id":"1610.05567","date":"2016-10-18","proceeding":null,"authors":["Rémi Cadène","Nicolas Thome","Matthieu Cord"],"abstract":"The goal of our research is to develop methods advancing automatic visual\nrecognition. In order to predict the unique or multiple labels associated to an\nimage, we study different kind of Deep Neural Networks architectures and\nmethods for supervised features learning. We first draw up a state-of-the-art\nreview of the Convolutional Neural Networks aiming to understand the history\nbehind this family of statistical models, the limit of modern architectures and\nthe novel techniques currently used to train deep CNNs. The originality of our\nwork lies in our approach focusing on tasks with a low amount of data. We\nintroduce different models and techniques to achieve the best accuracy on\nseveral kind of datasets, such as a medium dataset of food recipes (100k\nimages) for building a web API, or a small dataset of satellite images (6,000)\nfor the DSG online challenge that we've won. We also draw up the\nstate-of-the-art in Weakly Supervised Learning, introducing different kind of\nCNNs able to localize regions of interest. Our last contribution is a\nframework, build on top of Torch7, for training and testing deep models on any\nvisual recognition tasks and on datasets of any scale.","url_abs":"http://arxiv.org/abs/1610.05567v1","url_pdf":"http://arxiv.org/pdf/1610.05567v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"masters-thesis-deep-learning-for-visual","repo_url":"https://github.com/Cadene/torchnet-deep6","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"deep-learning","task_name":"Deep Learning"},{"task_slug":"weakly-supervised-learning","task_name":"Weakly-supervised Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}