{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/deep-cross-modal-learning-for-caricature","title":"Deep Cross Modal Learning for Caricature Verification and Identification(CaVINet)","arxiv_id":"1807.11688","date":"2018-07-31","proceeding":null,"authors":["Jatin Garg","Skand Vishwanath Peri","Himanshu Tolani","Narayanan C. Krishnan"],"abstract":"Learning from different modalities is a challenging task. In this paper, we\nlook at the challenging problem of cross modal face verification and\nrecognition between caricature and visual image modalities. Caricature have\nexaggerations of facial features of a person. Due to the significant variations\nin the caricatures, building vision models for recognizing and verifying data\nfrom this modality is an extremely challenging task. Visual images with\nsignificantly lesser amount of distortions can act as a bridge for the analysis\nof caricature modality. We introduce a publicly available large\nCaricature-VIsual dataset [CaVI] with images from both the modalities that\ncaptures the rich variations in the caricature of an identity. This paper\npresents the first cross modal architecture that handles extreme distortions of\ncaricatures using a deep learning network that learns similar representations\nacross the modalities. We use two convolutional networks along with\ntransformations that are subjected to orthogonality constraints to capture the\nshared and modality specific representations. In contrast to prior research,\nour approach neither depends on manually extracted facial landmarks for\nlearning the representations, nor on the identities of the person for\nperforming verification. The learned shared representation achieves 91%\naccuracy for verifying unseen images and 75% accuracy on unseen identities.\nFurther, recognizing the identity in the image by knowledge transfer using a\ncombination of shared and modality specific representations, resulted in an\nunprecedented performance of 85% rank-1 accuracy for caricatures and 95% rank-1\naccuracy for visual images.","url_abs":"http://arxiv.org/abs/1807.11688v1","url_pdf":"http://arxiv.org/pdf/1807.11688v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"deep-cross-modal-learning-for-caricature","repo_url":"https://github.com/lsaiml/CaVINet","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"caricature","task_name":"Caricature"},{"task_slug":"face-recognition","task_name":"Face Recognition"},{"task_slug":"face-verification","task_name":"Face Verification"},{"task_slug":"transfer-learning","task_name":"Transfer Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}