Diffusion for alignment, not generation
Text-guided denoising forms one skeleton–text latent space: best on all eight NTU zero-shot splits, by up to 13.05 points.
| Method | Venue | NTU-60 | NTU-120 | ||||||
|---|---|---|---|---|---|---|---|---|---|
| 55/5 | 48/12 | 40/20 | 30/30 | 110/10 | 96/24 | 80/40 | 60/60 | ||
| ReViSE | ICCV 2017 | 53.91 | 17.49 | 24.26 | 14.81 | 55.04 | 32.38 | 19.47 | 8.27 |
| JPoSE | ICCV 2019 | 64.82 | 28.75 | 20.05 | 12.39 | 51.93 | 32.44 | 13.71 | 7.65 |
| CADA-VAE | CVPR 2019 | 76.84 | 28.96 | 16.21 | 11.51 | 59.53 | 35.77 | 10.55 | 5.67 |
| SynSE | ICIP 2021 | 75.81 | 33.30 | 19.85 | 12.00 | 62.69 | 38.70 | 13.64 | 7.73 |
| SMIE | ACM MM 2023 | 77.98 | 40.18 | – | – | 65.74 | 45.30 | – | – |
| PURLS | CVPR 2024 | 79.23 | 40.99 | 31.05 | 23.52 | 71.95 | 52.01 | 28.38 | 19.63 |
| SA-DVAE | ECCV 2024 | 82.37 | 41.38 | – | – | 68.77 | 46.12 | – | – |
| STAR | ACM MM 2024 | 81.40 | 45.10 | – | – | 63.30 | 44.30 | – | – |
| TDSM (Ours) | – | 86.49 | 56.03 | 36.09 | 25.88 | 74.15 | 65.06 | 36.95 | 27.21 |
The SynSE splits (55/5, 48/12, 110/10, 96/24) are the standard settings; the PURLS splits (40/20, 30/30, 80/40, 60/60) hold out a larger share of unseen classes. Both use a Shift-GCN skeleton encoder, as in prior work.
Paper figures
Direct alignment versus diffusion-based alignment, accuracy across inference timesteps and noise draws (also when predicting zx), per-class results on two SMIE splits, the inference pipeline and the CrossDiT block.
Quantitative results
Best on NTU-60, NTU-120 and PKU-MMD under the SMIE benchmark, and on all eight Kinetics-200/400 splits with a single text prompt per action.
| Method | NTU-60 | NTU-120 | PKU-MMD |
|---|---|---|---|
| 55/5 | 110/10 | 46/5 | |
| ReViSE | 60.94 | 44.90 | 59.34 |
| JPoSE | 59.44 | 46.69 | 57.17 |
| CADA-VAE | 61.84 | 45.15 | 60.74 |
| SynSE | 64.19 | 47.28 | 53.85 |
| SMIE | 65.08 | 46.40 | 60.83 |
| SA-DVAE | 84.20 | 50.67 | 66.54 |
| STAR | 77.50 | – | 70.60 |
| TDSM (Ours) | 88.88 | 69.47 | 70.76 |
Average of three splits per dataset (each split: Table 10), with an ST-GCN skeleton encoder, as in prior work.
| Method | Kinetics-200 | |||
|---|---|---|---|---|
| 180/20 | 160/40 | 140/60 | 120/80 | |
| ReViSE | 24.95 | 13.28 | 8.14 | 6.23 |
| DeViSE | 22.22 | 12.32 | 7.97 | 5.65 |
| PURLS (1 text) | 25.96 | 15.85 | 10.23 | 7.77 |
| PURLS (7 text) | 32.22 | 22.56 | 12.01 | 11.75 |
| TDSM (1 text) | 38.18 | 24.43 | 15.28 | 13.09 |
| Method | Kinetics-400 | |||
|---|---|---|---|---|
| 360/40 | 320/80 | 300/100 | 280/120 | |
| ReViSE | 20.84 | 11.82 | 9.49 | 8.23 |
| DeViSE | 18.37 | 10.23 | 9.47 | 8.34 |
| PURLS (1 text) | 22.50 | 15.08 | 11.44 | 11.03 |
| PURLS (7 text) | 34.51 | 24.32 | 16.99 | 14.28 |
| TDSM (1 text) | 38.92 | 26.24 | 18.45 | 16.10 |
TDSM uses a single text prompt per action; PURLS is shown with one and with seven text prompts.
| Method | Modality | NTU-60 | NTU-120 | |||
|---|---|---|---|---|---|---|
| Text | RGB | 55/5 | 48/12 | 110/10 | 96/24 | |
| BSZSL | ✓ | ✓ | 83.04 | 52.96 | 77.69 | 56.12 |
| TDSM | ✓ | 86.49 | 56.03 | 74.15 | 65.06 | |
BSZSL uses RGB video in addition to text; without RGB input, TDSM is better on three of the four splits.
Ablations Tables 3–6
| ℒdiff | ℒTD | NTU-60 | NTU-120 | ||
|---|---|---|---|---|---|
| 55/5 | 48/12 | 110/10 | 96/24 | ||
| ✓ | 79.87 | 53.03 | 72.44 | 57.65 | |
| ✓ | 80.90 | 54.36 | 70.73 | 60.95 | |
| ✓ | ✓ | 86.49 | 56.03 | 74.15 | 65.06 |
| Global zg | Local zl | NTU-60 | NTU-120 | ||
|---|---|---|---|---|---|
| 55/5 | 48/12 | 110/10 | 96/24 | ||
| ✓ | 83.41 | 51.50 | 70.14 | 61.90 | |
| ✓ | 83.33 | 52.63 | 69.95 | 62.10 | |
| ✓ | ✓ | 86.49 | 56.03 | 74.15 | 65.06 |
| Total T | NTU-60 | NTU-120 | ||
|---|---|---|---|---|
| 55/5 | 48/12 | 110/10 | 96/24 | |
| 1 | 85.03 | 44.10 | 69.91 | 60.35 |
| 10 | 84.51 | 50.89 | 69.97 | 62.04 |
| 50 | 86.49 | 56.03 | 74.15 | 65.06 |
| 100 | 83.48 | 56.27 | 71.05 | 64.57 |
| 500 | 81.34 | 53.43 | 71.93 | 60.81 |
| Gaussian noise ϵ | NTU-60 | NTU-120 | ||
|---|---|---|---|---|
| 55/5 | 48/12 | 110/10 | 96/24 | |
| Fixed | 76.40 | 44.25 | 64.01 | 52.21 |
| Random | 86.49 | 56.03 | 74.15 | 65.06 |
The tinted row is the final TDSM setting in each ablation: both losses (λ = τ = 1.0), global and local text features, T = 50 and a new random noise at every training step.
Backbone and SMIE splits Tables 12 and 10
| Backbone | NTU-60 | NTU-120 | ||
|---|---|---|---|---|
| 55/5 | 48/12 | 110/10 | 96/24 | |
| U-Net | 82.40 | 51.12 | 70.03 | 59.77 |
| DiT (TDSM) | 86.49 | 56.03 | 74.15 | 65.06 |
The same framework with a U-Net denoiser instead of the DiT is lower on every split.
Comparison with previous ZSAR methods Table 11
| Method | Characteristics | Limitations |
|---|---|---|
| VAE-based | Reconstructs skeleton-text feature pairs via cross-reconstruction, recovering skeleton features from text and vice versa | Modality gap due to direct alignment |
| CL-based | Aligns skeleton and text features by minimizing feature distance through contrastive learning | |
| TDSM (Ours) | Denoises skeleton latents (i.e., estimates added noise in the forward diffusion) using reverse diffusion, conditioned on text embeddings, to naturally align both modalities in a unified latent space | Noise-sensitive performance |
The right label denoises best
- 01Text-conditioned denoising. Frozen skeleton (Shift-GCN or ST-GCN) and CLIP text encoders give the features; a Diffusion Transformer denoises the noisy skeleton feature conditioned on global and local text features, fusing both modalities in one latent space.
- 02Triplet diffusion loss. Next to ℒdiff, ℒTD = max(‖ϵ − ϵ̂p‖2 − ‖ϵ − ϵ̂n‖2 + τ, 0) rewards accurate denoising with the ground-truth prompt and suppresses it with a wrong one.
- 03One-step inference. Each candidate label’s prompt conditions a single denoising step at a fixed ttest = 25 and fixed noise ϵtest; the label whose predicted noise lies closest to ϵtest is the prediction.
BibTeX
@InProceedings{Do_2025_ICCV,
author={Do, Jeonghyeok and Kim, Munchurl},
title={Bridging the Skeleton-Text Modality Gap: Diffusion-Powered Modality Alignment for Zero-shot Skeleton-based Action Recognition},
booktitle={Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)},
month={October},
year={2025},
pages={12757-12768}
}