Papers › Multi-Semantic Fusion Model for Generalized Zero-Shot Skeleton-Based Action Recognition

Multi-Semantic Fusion Model for Generalized Zero-Shot Skeleton-Based Action Recognition

18 Sep 2023arXiv:2309.09592archive 2025-07-28

Ming-Zhe Li, Zhen Jia, Zhang Zhang, Zhanyu Ma, Liang Wang

Generalized zero-shot skeleton-based action recognition (GZSSAR) is a new challenging problem in computer vision community, which requires models to recognize actions without any training samples. Previous studies only utilize the action labels of verb phrases as the semantic prototypes for learning the mapping from skeleton-based actions to a shared semantic space. However, the limited semantic information of action labels restricts the generalization ability of skeleton features for recognizing unseen actions. In order to solve this dilemma, we propose a multi-semantic fusion (MSF) model for improving the performance of GZSSAR, where two kinds of class-level textual descriptions (i.e., action descriptions and motion descriptions), are collected as auxiliary semantic information to enhance the learning efficacy of generalizable skeleton features. Specially, a pre-trained language encoder takes the action descriptions, motion descriptions and original class labels as inputs to obtain rich semantic features for each action class, while a skeleton encoder is implemented to extract skeleton features. Then, a variational autoencoder (VAE) based generative module is performed to learn a cross-modal alignment between skeleton and semantic features. Finally, a classification module is built to recognize the action categories of input samples, where a seen-unseen classification gate is adopted to predict whether the sample comes from seen action classes or not in GZSSAR. The superior performance in comparisons with previous models validates the effectiveness of the proposed MSF model on GZSSAR.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

EHZ9NIWI7/MSF-GZSSAR officialmentioned in paperpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Action RecognitionGeneralized Zero Shot skeletal action recognitionSkeleton Based Action RecognitionZero-shot skeleton-based action recognitioncross-modal alignment

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Generalized Zero Shot skeletal action recognition NTU RGB+D MSF-GZSSAR Harmonic Mean (12 unseen classes) 49.70 #1 of 4 Archive leaderboard report
Generalized Zero Shot skeletal action recognition NTU RGB+D MSF-GZSSAR Harmonic Mean (5 unseen classes) 68.83 #1 of 4 Archive leaderboard report
Generalized Zero Shot skeletal action recognition NTU RGB+D 120 MSF-GZSSAR Harmonic Mean (10 unseen classes) 57.40 #2 of 4 Archive leaderboard report
Generalized Zero Shot skeletal action recognition NTU RGB+D 120 MSF-GZSSAR Harmonic Mean (24 unseen classes) 52.40 #2 of 4 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections