{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/modality-distillation-with-multiple-stream","title":"Modality Distillation with Multiple Stream Networks for Action Recognition","arxiv_id":"1806.07110","date":"2018-06-19","proceeding":"ECCV 2018 9","authors":["Nuno Garcia","Pietro Morerio","Vittorio Murino"],"abstract":"Diverse input data modalities can provide complementary cues for several\ntasks, usually leading to more robust algorithms and better performance.\nHowever, while a (training) dataset could be accurately designed to include a\nvariety of sensory inputs, it is often the case that not all modalities could\nbe available in real life (testing) scenarios, where a model has to be\ndeployed. This raises the challenge of how to learn robust representations\nleveraging multimodal data in the training stage, while considering limitations\nat test time, such as noisy or missing modalities.\n  This paper presents a new approach for multimodal video action recognition,\ndeveloped within the unified frameworks of distillation and privileged\ninformation, named generalized distillation. Particularly, we consider the case\nof learning representations from depth and RGB videos, while relying on RGB\ndata only at test time. We propose a new approach to train an hallucination\nnetwork that learns to distill depth features through multiplicative\nconnections of spatiotemporal representations, leveraging soft labels and hard\nlabels, as well as distance between feature maps. We report state-of-the-art\nresults on video action classification on the largest multimodal dataset\navailable for this task, the NTU RGB+D. Code available at\nhttps://github.com/ncgarcia/modality-distillation .","url_abs":"http://arxiv.org/abs/1806.07110v2","url_pdf":"http://arxiv.org/pdf/1806.07110v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"modality-distillation-with-multiple-stream","repo_url":"https://github.com/ncgarcia/modality-distillation","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"action-classification","task_name":"Action Classification"},{"task_slug":"action-detection","task_name":"Action Detection"},{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"hallucination","task_name":"Hallucination"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1806.07110","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}