{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/disentangled-variational-representation-for","title":"Disentangled Variational Representation for Heterogeneous Face Recognition","arxiv_id":"1809.01936","date":"2018-09-06","proceeding":null,"authors":["Xiang Wu","Huaibo Huang","Vishal M. Patel","Ran He","Zhenan Sun"],"abstract":"Visible (VIS) to near infrared (NIR) face matching is a challenging problem\ndue to the significant domain discrepancy between the domains and a lack of\nsufficient data for training cross-modal matching algorithms. Existing\napproaches attempt to tackle this problem by either synthesizing visible faces\nfrom NIR faces, extracting domain-invariant features from these modalities, or\nprojecting heterogeneous data onto a common latent space for cross-modal\nmatching. In this paper, we take a different approach in which we make use of\nthe Disentangled Variational Representation (DVR) for cross-modal matching.\nFirst, we model a face representation with an intrinsic identity information\nand its within-person variations. By exploring the disentangled latent variable\nspace, a variational lower bound is employed to optimize the approximate\nposterior for NIR and VIS representations. Second, aiming at obtaining more\ncompact and discriminative disentangled latent space, we impose a minimization\nof the identity information for the same subject and a relaxed correlation\nalignment constraint between the NIR and VIS modality variations. An\nalternative optimization scheme is proposed for the disentangled variational\nrepresentation part and the heterogeneous face recognition network part. The\nmutual promotion between these two parts effectively reduces the NIR and VIS\ndomain discrepancy and alleviates over-fitting. Extensive experiments on three\nchallenging NIR-VIS heterogeneous face recognition databases demonstrate that\nthe proposed method achieves significant improvements over the state-of-the-art\nmethods.","url_abs":"http://arxiv.org/abs/1809.01936v3","url_pdf":"http://arxiv.org/pdf/1809.01936v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"face-recognition","task_name":"Face Recognition"},{"task_slug":"heterogeneous-face-recognition","task_name":"Heterogeneous Face Recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/face-verification-on-buaa-visnir","task":"Face Verification","dataset":"BUAA-VisNir","model":"DVR Wu et al. (2019)","rank_in_archive_order":2,"of":3,"metrics":{"TAR @ FAR=0.001":"96.9","TAR @ FAR=0.01":"98.5"},"uses_additional_data":false},{"leaderboard":"/sota/face-verification-on-casia-nir-vis-20","task":"Face Verification","dataset":"CASIA NIR-VIS 2.0","model":"DVR Wu et al. (2019)","rank_in_archive_order":2,"of":3,"metrics":{"TAR @ FAR=0.001":"99.6"},"uses_additional_data":false},{"leaderboard":"/sota/face-verification-on-oulu-casia-nir-vis","task":"Face Verification","dataset":"Oulu-CASIA NIR-VIS","model":"DVR Wu et al. (2019)","rank_in_archive_order":2,"of":3,"metrics":{"TAR @ FAR=0.001":"84.9","TAR @ FAR=0.01":"97.2"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1809.01936","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}