Papers › TDSM: Triplet Diffusion for Skeleton-Text Matching in Zero-Shot Action Recognition

TDSM: Triplet Diffusion for Skeleton-Text Matching in Zero-Shot Action Recognition

16 Nov 2024arXiv:2411.10745archive 2025-07-28

Jeonghyeok Do, Munchurl Kim

We firstly present a diffusion-based action recognition with zero-shot learning for skeleton inputs. In zero-shot skeleton-based action recognition, aligning skeleton features with the text features of action labels is essential for accurately predicting unseen actions. Previous methods focus on direct alignment between skeleton and text latent spaces, but the modality gaps between these spaces hinder robust generalization learning. Motivated from the remarkable performance of text-to-image diffusion models, we leverage their alignment capabilities between different modalities mostly by focusing on the training process during reverse diffusion rather than using their generative power. Based on this, our framework is designed as a Triplet Diffusion for Skeleton-Text Matching (TDSM) method which aligns skeleton features with text prompts through reverse diffusion, embedding the prompts into the unified skeleton-text latent space to achieve robust matching. To enhance discriminative power, we introduce a novel triplet diffusion (TD) loss that encourages our TDSM to correct skeleton-text matches while pushing apart incorrect ones. Our TDSM significantly outperforms the very recent state-of-the-art methods with large margins of 2.36%-point to 13.05%-point, demonstrating superior accuracy and scalability in zero-shot settings through effective skeleton-text matching.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

KAIST-VICLab/TDSM officialmentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Action RecognitionSkeleton Based Action RecognitionText MatchingZero Shot Skeletal Action RecognitionZero-Shot Action RecognitionZero-Shot Learning

1 archive task tag without a task page not shown.

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Zero Shot Skeletal Action Recognition NTU RGB+D TDSM Accuracy (12 unseen classes) 56.03 #1 of 9 Archive leaderboard report
Zero Shot Skeletal Action Recognition NTU RGB+D TDSM Accuracy (5 unseen classes) 86.49 #1 of 9 Archive leaderboard report
Zero Shot Skeletal Action Recognition NTU RGB+D TDSM Random Split Accuracy 88.88 #1 of 9 Archive leaderboard report
Zero Shot Skeletal Action Recognition NTU RGB+D 120 TDSM Accuracy (10 unseen classes) 74.15 #2 of 9 Archive leaderboard report
Zero Shot Skeletal Action Recognition NTU RGB+D 120 TDSM Accuracy (24 unseen classes) 65.06 #2 of 9 Archive leaderboard report
Zero Shot Skeletal Action Recognition NTU RGB+D 120 TDSM Random Split Accuracy 69.47 #2 of 9 Archive leaderboard report
Zero Shot Skeletal Action Recognition PKU-MMD TDSM Random Split Accuracy 70.76 #2 of 7 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

DiffusionLatent Diffusion ModelTransformer

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections