{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/deep-binary-reconstruction-for-cross-modal","title":"Deep Binary Reconstruction for Cross-modal Hashing","arxiv_id":"1708.05127","date":"2017-08-17","proceeding":null,"authors":["Xuelong. Li","Di Hu","Feiping Nie"],"abstract":"With the increasing demand of massive multimodal data storage and\norganization, cross-modal retrieval based on hashing technique has drawn much\nattention nowadays. It takes the binary codes of one modality as the query to\nretrieve the relevant hashing codes of another modality. However, the existing\nbinary constraint makes it difficult to find the optimal cross-modal hashing\nfunction. Most approaches choose to relax the constraint and perform\nthresholding strategy on the real-value representation instead of directly\nsolving the original objective. In this paper, we first provide a concrete\nanalysis about the effectiveness of multimodal networks in preserving the\ninter- and intra-modal consistency. Based on the analysis, we provide a\nso-called Deep Binary Reconstruction (DBRC) network that can directly learn the\nbinary hashing codes in an unsupervised fashion. The superiority comes from a\nproposed simple but efficient activation function, named as Adaptive Tanh\n(ATanh). The ATanh function can adaptively learn the binary codes and be\ntrained via back-propagation. Extensive experiments on three benchmark datasets\ndemonstrate that DBRC outperforms several state-of-the-art methods in both\nimage2text and text2image retrieval task.","url_abs":"http://arxiv.org/abs/1708.05127v2","url_pdf":"http://arxiv.org/pdf/1708.05127v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"deep-binary-reconstruction-for-cross-modal","repo_url":"https://github.com/yolo2233/cross-modal-hasing-playground","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"cross-modal-retrieval","task_name":"Cross-Modal Retrieval"},{"task_slug":"retrieval","task_name":"Retrieval"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1708.05127","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}