Papers › HalluciNet-ing Spatiotemporal Representations Using a 2D-CNN

HalluciNet-ing Spatiotemporal Representations Using a 2D-CNN

10 Dec 2019arXiv:1912.04430archive 2025-07-28

Paritosh Parmar, Brendan Morris

Spatiotemporal representations learned using 3D convolutional neural networks (CNN) are currently used in state-of-the-art approaches for action related tasks. However, 3D-CNN are notorious for being memory and compute resource intensive as compared with more simple 2D-CNN architectures. We propose to hallucinate spatiotemporal representations from a 3D-CNN teacher with a 2D-CNN student. By requiring the 2D-CNN to predict the future and intuit upcoming activity, it is encouraged to gain a deeper understanding of actions and how they evolve. The hallucination task is treated as an auxiliary task, which can be used with any other action related task in a multitask learning setting. Thorough experimental evaluation shows that the hallucination task indeed helps improve performance on action recognition, action quality assessment, and dynamic scene recognition tasks. From a practical standpoint, being able to hallucinate spatiotemporal representations without an actual 3D-CNN can enable deployment in resource-constrained scenarios, such as with limited computing power and/or lower bandwidth. Codebase is available here: https://github.com/ParitoshParmar/HalluciNet.

PaperPDFCode

Code

ParitoshParmar/HalluciNet officialmentioned in papermentioned on GitHubpytorch report
ParitoshParmar/HalluciNet--PyTorch mentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Action AnticipationAction ClassificationAction Quality AssessmentAction RecognitionAction Recognition In Still ImagesFine-grained Action RecognitionHallucinationMulti-Task LearningScene Recognition

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Action Recognition UCF101 HalluciNet (ResNet-50) 3-fold Accuracy 79.83 #82 of 91 Archive leaderboard report
Scene Recognition YUP++ HalluciNet (ResNet-50) Accuracy (%) 84.44 #5 of 5 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Introduced by this paper: HalluciNet

HalluciNet

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections