{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/hallucinet-ing-spatiotemporal-representations","title":"HalluciNet-ing Spatiotemporal Representations Using a 2D-CNN","arxiv_id":"1912.04430","date":"2019-12-10","proceeding":null,"authors":["Paritosh Parmar","Brendan Morris"],"abstract":"Spatiotemporal representations learned using 3D convolutional neural networks (CNN) are currently used in state-of-the-art approaches for action related tasks. However, 3D-CNN are notorious for being memory and compute resource intensive as compared with more simple 2D-CNN architectures. We propose to hallucinate spatiotemporal representations from a 3D-CNN teacher with a 2D-CNN student. By requiring the 2D-CNN to predict the future and intuit upcoming activity, it is encouraged to gain a deeper understanding of actions and how they evolve. The hallucination task is treated as an auxiliary task, which can be used with any other action related task in a multitask learning setting. Thorough experimental evaluation shows that the hallucination task indeed helps improve performance on action recognition, action quality assessment, and dynamic scene recognition tasks. From a practical standpoint, being able to hallucinate spatiotemporal representations without an actual 3D-CNN can enable deployment in resource-constrained scenarios, such as with limited computing power and/or lower bandwidth. Codebase is available here: https://github.com/ParitoshParmar/HalluciNet.","url_abs":"https://arxiv.org/abs/1912.04430v3","url_pdf":"https://arxiv.org/pdf/1912.04430v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"hallucinet-ing-spatiotemporal-representations","repo_url":"https://github.com/ParitoshParmar/HalluciNet","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"hallucinet-ing-spatiotemporal-representations","repo_url":"https://github.com/ParitoshParmar/HalluciNet--PyTorch","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"action-anticipation","task_name":"Action Anticipation"},{"task_slug":"action-classification","task_name":"Action Classification"},{"task_slug":"action-quality-assessment","task_name":"Action Quality Assessment"},{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"action-recognition-in-still-images","task_name":"Action Recognition In Still Images"},{"task_slug":"fine-grained-action-recognition","task_name":"Fine-grained Action Recognition"},{"task_slug":"hallucination","task_name":"Hallucination"},{"task_slug":"multi-task-learning","task_name":"Multi-Task Learning"},{"task_slug":"scene-recognition","task_name":"Scene Recognition"}],"methods":[{"method_slug":"hallucinet","method_name":"HalluciNet"}],"datasets_introduced":[],"methods_introduced":[{"slug":"hallucinet","name":"HalluciNet","full_name":"Approximating Spatiotemporal Representations Using a 2DCNN"}],"results":[{"leaderboard":"/sota/action-recognition-in-videos-on-ucf101","task":"Action Recognition","dataset":"UCF101","model":"HalluciNet (ResNet-50)","rank_in_archive_order":82,"of":91,"metrics":{"3-fold Accuracy":"79.83"},"uses_additional_data":false},{"leaderboard":"/sota/scene-recognition-on-yup","task":"Scene Recognition","dataset":"YUP++","model":"HalluciNet (ResNet-50)","rank_in_archive_order":5,"of":5,"metrics":{"Accuracy (%)":"84.44"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}