{"url":"/method/dolg","slug":"dolg","name":"DOLG","full_name":"Deep Orthogonal Fusion of Local and Global Features","full_name_withheld":false,"description_markdown":"Image Retrieval is a fundamental task of obtaining images similar to the query one from a database. A common image retrieval practice is to firstly retrieve candidate images via similarity search using global image features and then re-rank the candidates by leveraging their\r\nlocal features. Previous learning-based studies mainly focus on either global or local image representation learning\r\nto tackle the retrieval task. In this paper, we abandon the\r\ntwo-stage paradigm and seek to design an effective singlestage solution by integrating local and global information\r\ninside images into compact image representations. Specifically, we propose a Deep Orthogonal Local and Global\r\n(DOLG) information fusion framework for end-to-end image retrieval. It attentively extracts representative local information with multi-atrous convolutions and self-attention\r\nat first. Components orthogonal to the global image representation are then extracted from the local information.\r\nAt last, the orthogonal components are concatenated with\r\nthe global representation as a complementary, and then aggregation is performed to generate the final representation.\r\nThe whole framework is end-to-end differentiable and can\r\nbe trained with image-level labels. Extensive experimental\r\nresults validate the effectiveness of our solution and show\r\nthat our model achieves state-of-the-art image retrieval performances on Revisited Oxford and Paris datasets.","description_state":"present","introduced_year":null,"introduced_by":{"title":"DOLG: Single-Stage Image Retrieval with Deep Orthogonal Fusion of Local and Global Features","paper":"/paper/dolg-single-stage-image-retrieval-with-deep","first_author":"Min Yang","n_authors":8,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/dolg-single-stage-image-retrieval-with-deep"},"source":{"url":"https://arxiv.org/abs/2108.02927v2","title":"DOLG: Single-Stage Image Retrieval with Deep Orthogonal Fusion of Local and Global Features","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Computer Vision","area_id":"computer-vision","collection":"Image Retrieval Models","url":"/methods/category/image-retrieval-models","pwc_aliases":[]}],"n_papers_tagged":1,"archive_num_papers":1,"papers_newest_first":[{"paper":"/paper/dolg-single-stage-image-retrieval-with-deep","title":"DOLG: Single-Stage Image Retrieval with Deep Orthogonal Fusion of Local and Global Features","date":"2021-08-06","arxiv_id":"2108.02927","n_code_links":5,"syntology":{"ran":10,"of":23,"unverified":13,"pointer_only":0}}],"papers_shown":1,"tasks":[{"task":"/task/image-retrieval","name":"Image Retrieval","papers":1},{"task":"/task/representation-learning","name":"Representation Learning","papers":1},{"task":"/task/retrieval","name":"Retrieval","papers":1}],"tasks_shown":3,"n_tasks":3,"usage_by_year":[{"year":"2021","papers":1}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/dolg"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}