More research

Triplet Diffusion for Skeleton-Text Matching Bridging the Skeleton-Text Modality Gap: Diffusion-Powered Modality Alignment for Zero-shot Skeleton-based Action Recognition

Jeonghyeok Do Munchurl Kim†

Korea Advanced Institute of Science and Technology (KAIST), South Korea

† Corresponding authorehwjdgur0913@kaist.ac.krmkimee@kaist.ac.kr

ICCV 2025

Abstract

In zero-shot skeleton-based action recognition (ZSAR), aligning skeleton features with the text features of action labels is essential for accurately predicting unseen actions. ZSAR faces a fundamental challenge in bridging the modality gap between the two-kind features, which severely limits generalization to unseen actions. Previous methods focus on direct alignment between skeleton and text latent spaces, but the modality gaps between these spaces hinder robust generalization learning. Motivated by the success of diffusion models in multi-modal alignment (e.g., text-to-image, text-to-video), we firstly present a diffusion-based skeleton-text alignment framework for ZSAR. Our approach, Triplet Diffusion for Skeleton-Text Matching (TDSM), focuses on cross-alignment power of diffusion models rather than their generative capability. Specifically, TDSM aligns skeleton features with text prompts by incorporating text features into the reverse diffusion process, where skeleton features are denoised under text guidance, forming a unified skeleton-text latent space for robust matching. To enhance discriminative power, we introduce a triplet diffusion (TD) loss that encourages our TDSM to correct skeleton-text matches while pushing them apart for different action classes. Our TDSM significantly outperforms very recent state-of-the-art methods with significantly large margins of 2.36%-point to 13.05%-point, demonstrating superior accuracy and scalability in zero-shot settings through effective skeleton-text matching.

Diffusion for alignment, not generation

Text-guided denoising forms one skeleton–text latent space: best on all eight NTU zero-shot splits, by up to 13.05 points.

Table 1. Top-1 accuracy results of various zero-shot skeleton-based action recognition (ZSAR) methods evaluated on the SynSE and PURLS benchmarks for the NTU-60 and NTU-120 datasets. Each split is denoted as X/Y, where X represents the number of seen classes and Y the number of unseen classes. For our TDSM framework, the reported accuracy is the average value obtained from 10 trials, each with different Gaussian noise.
Bold bestUnderline second bestTop-1 accuracy (%)
Scroll for more columns
MethodVenueNTU-60NTU-120
55/548/1240/2030/30110/1096/2480/4060/60
ReViSEICCV 201753.9117.4924.2614.8155.0432.3819.478.27
JPoSEICCV 201964.8228.7520.0512.3951.9332.4413.717.65
CADA-VAECVPR 201976.8428.9616.2111.5159.5335.7710.555.67
SynSEICIP 202175.8133.3019.8512.0062.6938.7013.647.73
SMIEACM MM 202377.9840.18––65.7445.30––
PURLSCVPR 202479.2340.9931.0523.5271.9552.0128.3819.63
SA-DVAEECCV 202482.3741.38––68.7746.12––
STARACM MM 202481.4045.10––63.3044.30––
TDSM (Ours)–86.4956.0336.0925.8874.1565.0636.9527.21

The SynSE splits (55/5, 48/12, 110/10, 96/24) are the standard settings; the PURLS splits (40/20, 30/30, 80/40, 60/60) hold out a larger share of unseen classes. Both use a Shift-GCN skeleton encoder, as in prior work.

Paper figures

Direct alignment versus diffusion-based alignment, accuracy across inference timesteps and noise draws (also when predicting zx), per-class results on two SMIE splits, the inference pipeline and the CrossDiT block.

Quantitative results

Best on NTU-60, NTU-120 and PKU-MMD under the SMIE benchmark, and on all eight Kinetics-200/400 splits with a single text prompt per action.

Table 2. Top-1 accuracy results of various ZSAR methods evaluated on the NTU-60, NTU-120, and PKU-MMD datasets under the SMIE benchmark. The reported values are the average performance across three splits.
Bold bestUnderline second bestTop-1 accuracy (%)
Scroll for more columns
MethodNTU-60NTU-120PKU-MMD
55/5110/1046/5
ReViSE60.9444.9059.34
JPoSE59.4446.6957.17
CADA-VAE61.8445.1560.74
SynSE64.1947.2853.85
SMIE65.0846.4060.83
SA-DVAE84.2050.6766.54
STAR77.50–70.60
TDSM (Ours)88.8869.4770.76

Average of three splits per dataset (each split: Table 10), with an ST-GCN skeleton encoder, as in prior work.

Ablations Tables 3–6
Table 3. Ablation study on loss function configurations. The results compare models trained with only the diffusion loss ℒdiff, only the triplet diffusion loss ℒTD, and their combination.
Top-1 accuracy (%)
Scroll for more columns
ℒdiffℒTDNTU-60NTU-120
55/548/12110/1096/24
✓79.8753.0372.4457.65
✓80.9054.3670.7360.95
✓✓86.4956.0374.1565.06
Table 4. Ablation study on text feature types. The results compare models trained with only zg, only zl, and their combination.
Top-1 accuracy (%)
Scroll for more columns
Global
zg
Local
zl
NTU-60NTU-120
55/548/12110/1096/24
✓83.4151.5070.1461.90
✓83.3352.6369.9562.10
✓✓86.4956.0374.1565.06
Table 5. Ablation study on the impact of total timesteps T in the training of the diffusion process.
Top-1 accuracy (%)
Scroll for more columns
Total TNTU-60NTU-120
55/548/12110/1096/24
185.0344.1069.9160.35
1084.5150.8969.9762.04
5086.4956.0374.1565.06
10083.4856.2771.0564.57
50081.3453.4371.9360.81
Table 6. Ablation study on the effect of noise ϵ during training.
Top-1 accuracy (%)
Scroll for more columns
Gaussian
noise ϵ
NTU-60NTU-120
55/548/12110/1096/24
Fixed76.4044.2564.0152.21
Random86.4956.0374.1565.06

The tinted row is the final TDSM setting in each ablation: both losses (λ = τ = 1.0), global and local text features, T = 50 and a new random noise at every training step.

Backbone and SMIE splits Tables 12 and 10
Table 12. Diffusion backbone: U-Net versus DiT (TDSM).
Top-1 accuracy (%)
Scroll for more columns
BackboneNTU-60NTU-120
55/548/12110/1096/24
U-Net82.4051.1270.0359.77
DiT (TDSM)86.4956.0374.1565.06

The same framework with a U-Net denoiser instead of the DiT is lower on every split.

Table 10. Top-1 accuracy results of our TDSM evaluated on the NTU-60, NTU-120, and PKU-MMD datasets under the SMIE benchmark.
Top-1 accuracy (%)
Scroll for more columns
TDSM (Ours)NTU-60NTU-120PKU-MMD
55/5110/1046/5
Split 187.9774.4557.40 (Fig. 7)
Split 296.06 (Fig. 6)63.9176.92
Split 382.6070.0477.97
Average88.8869.4770.76

The unseen actions of NTU-60 split 2 have distinct motion patterns (Fig. 6); four of the five in PKU-MMD split 1 share upward hand movements (Fig. 7).

Comparison with previous ZSAR methods Table 11
Table 11. Comparison of our TDSM with existing ZSAR methods.
Scroll for more columns
MethodCharacteristicsLimitations
VAE-basedReconstructs skeleton-text feature pairs via cross-reconstruction, recovering skeleton features from text and vice versaModality gap due to direct alignment
CL-basedAligns skeleton and text features by minimizing feature distance through contrastive learning
TDSM (Ours)Denoises skeleton latents (i.e., estimates added noise in the forward diffusion) using reverse diffusion, conditioned on text embeddings, to naturally align both modalities in a unified latent spaceNoise-sensitive performance

The right label denoises best

TDSM training. A seen skeleton sequence labelled Throw is encoded by a frozen skeleton encoder; noise epsilon is added at a random timestep t to give the noisy feature z_x,t. The ground-truth label Throw and a wrong label Pickup are turned into prompts and encoded by a frozen text encoder into global and local text features. A Diffusion Transformer with z_x, t, z_g and z_l embeddings, CrossDiT blocks, layer norm and a linear layer predicts the noise twice with shared weights: epsilon-hat_p for the positive prompt and epsilon-hat_n for the negative one. The diffusion loss is the distance between epsilon and epsilon-hat_p; the triplet diffusion loss is max of that distance minus the distance to epsilon-hat_n plus tau, and 0; the total loss is L_diff plus lambda L_TD.
Training framework of TDSM. The ground-truth prompt and a wrong-label prompt condition the same Diffusion Transformer; ℒdiff and ℒTD reward accurate denoising for the correct pair only.
  1. 01Text-conditioned denoising. Frozen skeleton (Shift-GCN or ST-GCN) and CLIP text encoders give the features; a Diffusion Transformer denoises the noisy skeleton feature conditioned on global and local text features, fusing both modalities in one latent space.
  2. 02Triplet diffusion loss. Next to ℒdiff, ℒTD = max(‖ϵ − ϵ̂p‖2 − ‖ϵ − ϵ̂n‖2 + τ, 0) rewards accurate denoising with the ground-truth prompt and suppresses it with a wrong one.
  3. 03One-step inference. Each candidate label’s prompt conditions a single denoising step at a fixed ttest = 25 and fixed noise ϵtest; the label whose predicted noise lies closest to ϵtest is the prediction.

BibTeX

@InProceedings{Do_2025_ICCV,
  author={Do, Jeonghyeok and Kim, Munchurl},
  title={Bridging the Skeleton-Text Modality Gap: Diffusion-Powered Modality Alignment for Zero-shot Skeleton-based Action Recognition},
  booktitle={Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)},
  month={October},
  year={2025},
  pages={12757-12768}
}