Papers › MMNet: A Model-Based Multimodal Network for Human Action Recognition in RGB-D Videos

MMNet: A Model-Based Multimodal Network for Human Action Recognition in RGB-D Videos

26 May 2022IEEE Transactions on Pattern Analysis and Machine Intelligence 2022 5archive 2025-07-28

Bruce X.B. Yu, Yan Liu, Xiang Zhang, Sheng-hua Zhong, Keith C.C. Chan

Human action recognition (HAR) in RGB-D videos has been widely investigated since the release of affordable depth sensors. Currently, unimodal approaches (e.g., skeleton-based and RGB video-based) have realized substantial improvements with increasingly larger datasets. However, multimodal methods specifically with model-level fusion have seldom been investigated. In this paper, we propose a model-based multimodal network (MMNet) that fuses skeleton and RGB modalities via a model-based approach. The objective of our method is to improve ensemble recognition accuracy by effectively applying mutually complementary information from different data modalities. For the model-based fusion scheme, we use a spatiotemporal graph convolution network for the skeleton modality to learn attention weights that will be transferred to the network of the RGB modality. Extensive experiments are conducted on five benchmark datasets: NTU RGB+D 60, NTU RGB+D 120, PKU-MMD, Northwestern-UCLA Multiview, and Toyota Smarthome. Upon aggregating the results of multiple modalities, our method is found to outperform state-of-the-art approaches on six evaluation protocols of the five datasets; thus, the proposed MMNet can effectively capture mutually complementary features in different RGB-D video modalities and provide more discriminative features for HAR. We also tested our MMNet on an RGB video dataset Kinetics 400 that contains more outdoor actions, which shows consistent results with those of RGB-D video datasets.

PaperPDFCode

Code

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Action ClassificationAction RecognitionAction Recognition In VideosSkeleton Based Action RecognitionTemporal Action Localization

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Action Classification Toyota Smarthome dataset MMNet CS 70.1 #3 of 13 Archive leaderboard report
Action Recognition NTU RGB+D MMNet (RGB + Pose) Accuracy (CS) 96.0 #6 of 28 Archive leaderboard report
Action Recognition NTU RGB+D MMNet (RGB + Pose) Accuracy (CV) 98.8 #6 of 28 Archive leaderboard report
Action Recognition NTU RGB+D 120 MMNet (RGB + Pose) Accuracy (Cross-Setup) 94.4 #4 of 21 Archive leaderboard report
Action Recognition NTU RGB+D 120 MMNet (RGB + Pose) Accuracy (Cross-Subject) 92.9 #4 of 21 Archive leaderboard report
Action Recognition In Videos PKU-MMD MMNet X-Sub 97.4 #2 of 5 Archive leaderboard report
Action Recognition In Videos PKU-MMD MMNet X-View 98.6 #2 of 5 Archive leaderboard report
Skeleton Based Action Recognition N-UCLA MMNet (RGB + Pose) Accuracy 93.7 #18 of 25 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

EfficientNet

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections