{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-local-image-descriptors-with-deep","title":"Learning Local Image Descriptors with Deep Siamese and Triplet Convolutional Networks by Minimising Global Loss Functions","arxiv_id":"1512.09272","date":"2015-12-31","proceeding":"CVPR 2016 6","authors":["Vijay Kumar B G","Gustavo Carneiro","Ian Reid"],"abstract":"Recent innovations in training deep convolutional neural network (ConvNet)\nmodels have motivated the design of new methods to automatically learn local\nimage descriptors. The latest deep ConvNets proposed for this task consist of a\nsiamese network that is trained by penalising misclassification of pairs of\nlocal image patches. Current results from machine learning show that replacing\nthis siamese by a triplet network can improve the classification accuracy in\nseveral problems, but this has yet to be demonstrated for local image\ndescriptor learning. Moreover, current siamese and triplet networks have been\ntrained with stochastic gradient descent that computes the gradient from\nindividual pairs or triplets of local image patches, which can make them prone\nto overfitting. In this paper, we first propose the use of triplet networks for\nthe problem of local image descriptor learning. Furthermore, we also propose\nthe use of a global loss that minimises the overall classification error in the\ntraining set, which can improve the generalisation capability of the model.\nUsing the UBC benchmark dataset for comparing local image descriptors, we show\nthat the triplet network produces a more accurate embedding than the siamese\nnetwork in terms of the UBC dataset errors. Moreover, we also demonstrate that\na combination of the triplet and global losses produces the best embedding in\nthe field, using this triplet network. Finally, we also show that the use of\nthe central-surround siamese network trained with the global loss produces the\nbest result of the field on the UBC dataset. Pre-trained models are available\nonline at https://github.com/vijaykbg/deep-patchmatch","url_abs":"http://arxiv.org/abs/1512.09272v2","url_pdf":"http://arxiv.org/pdf/1512.09272v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-local-image-descriptors-with-deep","repo_url":"https://github.com/vijaykbg/deep-patchmatch","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok","spdx":"Apache-2.0"}},{"paper_slug":"learning-local-image-descriptors-with-deep","repo_url":"https://github.com/sk1712/gcn_metric_learning","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"classification","task_name":"General Classification"},{"task_slug":null,"task_name":"Triplet"}],"methods":[{"method_slug":"siamese-network","method_name":"Siamese Network"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1512.09272","atlas_url":"https://app.syntology.ai/?focus=1512.09272","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}