{"url":"/method/delg","slug":"delg","name":"DELG","full_name":"DELG","full_name_withheld":false,"description_markdown":"**DELG** is a convolutional neural network for image retrieval that combines generalized mean pooling for global features and attentive selection for local features. The entire network can be learned end-to-end by carefully balancing the gradient flow between two heads – requiring only image-level labels. This allows for efficient inference by extracting an image’s global feature, detected keypoints and local descriptors within a single model.\r\n\r\nThe model is enabled by leveraging hierarchical image representations that arise in [CNNs](https://paperswithcode.com/methods/category/convolutional-neural-networks), which are coupled to [generalized mean pooling](https://paperswithcode.com/method/generalized-mean-pooling) and attentive local feature detection. Secondly, a convolutional autoencoder module is adopted that can successfully learn low-dimensional local descriptors. This can be readily integrated into the unified model, and avoids the need of post-processing learning steps, such as [PCA](https://paperswithcode.com/method/pca), that are commonly used. Finally, a procedure is used that enables end-to-end training of the proposed model using only image-level supervision. This requires carefully controlling the gradient flow between the global and local network heads during backpropagation, to avoid disrupting the desired representations.","description_state":"present","introduced_year":null,"introduced_by":{"title":"Unifying Deep Local and Global Features for Image Search","paper":"/paper/unifying-deep-local-and-global-features-for","first_author":"Bingyi Cao","n_authors":3,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/unifying-deep-local-and-global-features-for"},"source":{"url":"https://arxiv.org/abs/2001.05027v4","title":"Unifying Deep Local and Global Features for Image Search","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Computer Vision","area_id":"computer-vision","collection":"Image Retrieval Models","url":"/methods/category/image-retrieval-models","pwc_aliases":[]},{"area":"Computer Vision","area_id":"computer-vision","collection":"Convolutional Neural Networks","url":"/methods/category/convolutional-neural-networks","pwc_aliases":[]}],"n_papers_tagged":2,"archive_num_papers":2,"papers_newest_first":[{"paper":null,"title":"Deep Learning Based Image Retrieval in the JPEG Compressed Domain","date":"2021-07-08","arxiv_id":"2107.03648","n_code_links":0,"syntology":null},{"paper":"/paper/unifying-deep-local-and-global-features-for","title":"Unifying Deep Local and Global Features for Image Search","date":"2020-01-14","arxiv_id":"2001.05027","n_code_links":5,"syntology":null}],"papers_shown":2,"tasks":[{"task":"/task/image-retrieval","name":"Image Retrieval","papers":2},{"task":"/task/retrieval","name":"Retrieval","papers":2},{"task":"/task/content-based-image-retrieval","name":"Content-Based Image Retrieval","papers":1},{"task":"/task/deep-learning","name":"Deep Learning","papers":1},{"task":"/task/dimensionality-reduction","name":"Dimensionality Reduction","papers":1}],"tasks_shown":5,"n_tasks":5,"usage_by_year":[{"year":"2020","papers":1},{"year":"2021","papers":1}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/delg"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}