{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/vitac-feature-sharing-between-vision-and","title":"ViTac: Feature Sharing between Vision and Tactile Sensing for Cloth Texture Recognition","arxiv_id":"1802.07490","date":"2018-02-21","proceeding":null,"authors":["Shan Luo","Wenzhen Yuan","Edward Adelson","Anthony G. Cohn","Raul Fuentes"],"abstract":"Vision and touch are two of the important sensing modalities for humans and\nthey offer complementary information for sensing the environment. Robots could\nalso benefit from such multi-modal sensing ability. In this paper, addressing\nfor the first time (to the best of our knowledge) texture recognition from\ntactile images and vision, we propose a new fusion method named Deep Maximum\nCovariance Analysis (DMCA) to learn a joint latent space for sharing features\nthrough vision and tactile sensing. The features of camera images and tactile\ndata acquired from a GelSight sensor are learned by deep neural networks. But\nthe learned features are of a high dimensionality and are redundant due to the\ndifferences between the two sensing modalities, which deteriorates the\nperception performance. To address this, the learned features are paired using\nmaximum covariance analysis. Results of the algorithm on a newly collected\ndataset of paired visual and tactile data relating to cloth textures show that\na good recognition performance of greater than 90\\% can be achieved by using\nthe proposed DMCA framework. In addition, we find that the perception\nperformance of either vision or tactile sensing can be improved by employing\nthe shared representation space, compared to learning from unimodal data.","url_abs":"http://arxiv.org/abs/1802.07490v2","url_pdf":"http://arxiv.org/pdf/1802.07490v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"vitac-feature-sharing-between-vision-and","repo_url":"https://github.com/jettdlee/vis_tac_cross_modal","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1802.07490","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}