Fewer tokens, higher accuracy
A compact 8 × 25 token grid: state-of-the-art accuracy with 7.89× less inference computation than dense-token MAEs.
| Method | Tokens | Inference (GFLOPs) | Training (GFLOPs) | |
|---|---|---|---|---|
| NT × NJ | Encoder | Encoder | Decoder | |
| SkeletonMAE | 30 × 25 | 28.32 | 1.97 | 17.70 |
| MAMP | 30 × 25 | 28.32 | 1.97 | 17.70 |
| S-JEPA | 30 × 25 | 28.32 | 1.97 | 17.70 |
| GFP | 30 × 25 | 28.32 | 1.97 | 1.57 |
| SLiM (Ours) | 8 × 25 | 3.59 | 3.59 | – |
Paper figures
Feature prediction needs neither dense tokens nor a reconstruction decoder; Semantic Tube Masking and Skeleton-Aware Augmentations prevent shortcut learning.
Quantitative results
Best on four of five linear-evaluation protocols, three of four semi-supervised settings and both retrieval protocols, using a single 3D joint stream.
| Method | Venue | NTU-60 | NTU-120 | PKU-MMD II | ||
|---|---|---|---|---|---|---|
| X-Sub | X-View | X-Sub | X-Set | X-Sub | ||
| Other skeleton SSL baselines | ||||||
| GL-Transformer | ECCV'22 | 76.3 | 83.8 | 66.0 | 68.7 | – |
| MacDiff | ECCV'24 | 86.4 | 91.0 | 79.4 | 80.2 | – |
| IGM | ECCV'24 | 86.2 | 91.2 | 80.0 | 81.4 | – |
| USDRL | AAAI'25 | 85.2 | 91.7 | 76.6 | 78.1 | 54.4 |
| HSP | CVPR'25 | 80.7 | 88.0 | 71.0 | 73.2 | 48.9 |
| Contrastive learning (CL) | ||||||
| CPM | ECCV'22 | 78.7 | 84.9 | 68.7 | 69.6 | 48.3 |
| CMD | ECCV'22 | 79.8 | 86.9 | 70.3 | 71.5 | 43.0 |
| AimCLR | AAAI'22 | 74.3 | 79.7 | 63.4 | 63.4 | – |
| RVTCLR | ICCV'23 | 74.7 | 79.1 | – | – | – |
| PTSL | AAAI'23 | 77.3 | 81.8 | 66.2 | 67.7 | 49.3 |
| HYSP | ICLR'23 | 78.2 | 82.6 | 61.8 | 64.6 | – |
| HaLP | CVPR'23 | 79.7 | 86.8 | 71.1 | 72.2 | 43.5 |
| ActCLR | CVPR'23 | 80.9 | 86.7 | 69.0 | 70.5 | – |
| Masked / predictive modeling | ||||||
| SkeletonMAE | ICMEW'23 | 74.8 | 77.7 | 72.5 | 73.5 | 36.1 |
| MAMP | ICCV'23 | 84.9 | 89.1 | 78.6 | 79.1 | 53.8 |
| S-JEPA | ECCV'24 | 85.3 | 89.8 | 79.6 | 79.9 | 53.5 |
| GFP | ICCV'25 | 85.9 | 92.0 | 79.1 | 80.3 | 56.2 |
| AMR | CVPR'26 | 87.4 | 92.3 | 81.1 | 81.9 | 60.3 |
| Unified MFP + CL | ||||||
| SLiM (Ours) | – | 87.9 | 93.2 | 81.2 | 83.6 | 59.7 |
Including the concurrent AMR, SLiM is best on four of five protocols; on PKU-MMD II it ranks second, 0.6 points below AMR. All results use the joint modality only, without bone or motion streams.
(a) Semi-supervised learning
| Method | X-Sub | X-View | ||
|---|---|---|---|---|
| 1% | 10% | 1% | 10% | |
| CPM | 56.7 | 73.0 | 57.5 | 77.1 |
| CMD | 50.6 | 75.4 | 53.0 | 80.2 |
| HYSP | – | 76.2 | – | 80.4 |
| HaLP | 46.6 | 72.6 | 48.7 | 77.1 |
| HiCo | 54.4 | 73.0 | 54.8 | 78.3 |
| USDRL | 57.3 | 80.2 | 60.7 | 84.0 |
| SkeletonMAE | 54.4 | 80.6 | 54.6 | 83.5 |
| MAMP | 66.0 | 88.0 | 68.7 | 91.5 |
| S-JEPA | 67.5 | 88.4 | 69.1 | 91.4 |
| GFP | 71.8 | 88.7 | 72.9 | 92.1 |
| SLiM (Ours) | 72.1 | 88.8 | 75.8 | 91.9 |
(b) Action retrieval
| Method | NTU-60 | |
|---|---|---|
| X-Sub | X-View | |
| LongT-GAN | 39.1 | 48.1 |
| P&C | 50.7 | 76.3 |
| ISC | 62.5 | 82.6 |
| HaLP | 65.8 | 83.6 |
| HiCo | 68.3 | 84.8 |
| SkeAttnCLR | 69.4 | 76.8 |
| UmURL | 71.3 | 88.3 |
| MAMP | 62.0 | 70.0 |
| GFP | 70.9 | 87.1 |
| SLiM (Ours) | 72.5 | 89.9 |
With 1% labels SLiM improves the previous best by 0.3 (X-Sub) and 2.9 (X-View) points. Retrieval uses the frozen features without training a classifier.
| Method | To PKU-II | |
|---|---|---|
| NTU-60 | NTU-120 | |
| LongT-GAN | 44.8 | – |
| ISC | 51.1 | 52.3 |
| CMD | 56.0 | 57.0 |
| MacDiff | 72.2 | 73.4 |
| IGM | 59.8 | – |
| SkeletonMAE | 58.4 | 61.0 |
| MAMP | 70.6 | 73.2 |
| S-JEPA | 71.4 | 74.2 |
| SLiM (Ours) | 72.7 | 75.3 |
Ablations Tables 4 and 5
| Objective | NTU-60 | |||
|---|---|---|---|---|
| ℒMFP | ℒCL | ℒGLCL | X-Sub | X-View |
| ✗ | ✓ | ✗ | 73.6 | 78.9 |
| ✗ | ✗ | ✓ | 75.3 | 80.1 |
| ✓ | ✗ | ✗ | 85.3 | 90.6 |
| ✓ | ✓ | ✗ | 86.2 | 91.3 |
| ✓ | ✗ | ✓ | 87.7 | 92.7 |
ℒMFP alone reaches 85.3% on X-Sub; adding a standard CL loss gives 86.2%, and ℒGLCL instead gives 87.7%. Ablation models are pre-trained for 100 epochs (final model: 150).
(a) Semantic Tube Masking
| Masking | NTU-60 | ||
|---|---|---|---|
| Global view | Local view | X-Sub | X-View |
| ✗ | ✓ | 75.3 | 80.1 |
| ✓ (Limited) | ✓ | 84.2 | 90.4 |
| ✓ | ✗ | 83.3 | 89.4 |
| ✓ | ✓ (Limited) | 85.7 | 89.9 |
| ✓ (90%) | ✓ (90%) | 86.8 | 92.1 |
| ✓ (0–50%) | ✓ (0–50%) | 84.5 | 89.6 |
| ✓ (50–90%) | ✓ (50–90%) | 87.7 | 92.7 |
(b) Skeleton-Aware Augmentations
| Rotation | Mirroring | Scaling | X-Sub | X-View |
|---|---|---|---|---|
| ✗ | ✓ | ✓ | 81.6 | 86.9 |
| ✓ (Limited) | ✓ | ✓ | 84.8 | 89.9 |
| ✓ | ✗ | ✓ | 85.3 | 90.6 |
| ✓ | ✓ (Limited) | ✓ | 85.9 | 91.0 |
| ✓ | ✓ | ✗ | 84.4 | 90.0 |
| ✓ | ✓ | ✓ (Limited) | 86.4 | 91.7 |
| ✓ | ✓ | ✓ | 87.7 | 92.7 |
Replacing STM with conventional joint masking in either role lowers X-Sub by up to 3.5 points (87.7 vs. 84.2); the stochastic 50–90% ratio beats a fixed 90% by 0.9 and 0–50% by 3.2 points. Limited is the conventional counterpart shown in Fig. 4 (a)–(d).
Token budget and cost Tables 7 and 8
| Method | Clip T | Patch PT | Tokens NT × NJ | Inference GFLOPs | Avg. Acc. (%) |
|---|---|---|---|---|---|
| SkeletonMAE-8 | 64 | 8 | 8 × 25 | 3.59 | 70.50 |
| SkeletonMAE-15 | 120 | 8 | 15 × 25 | 8.36 | 74.00 |
| SkeletonMAE-30 | 120 | 4 | 30 × 25 | 28.32 | 76.25 |
| MAMP-8 | 64 | 8 | 8 × 25 | 3.59 | 80.90 |
| MAMP-15 | 120 | 8 | 15 × 25 | 8.36 | 85.50 |
| MAMP-30 | 120 | 4 | 30 × 25 | 28.32 | 87.00 |
| S-JEPA | 120 | 4 | 30 × 25 | 28.32 | 87.55 |
| GFP | 120 | 4 | 30 × 25 | 28.32 | 88.95 |
| SLiM-8 (Batch 384) | 64 | 8 | 8 × 25 | 3.59 | 89.60 |
| SLiM-8 (Batch 768) | 64 | 8 | 8 × 25 | 3.59 | 90.55 |
| SLiM-15 | 120 | 8 | 15 × 25 | 8.36 | 90.75 |
| SLiM-30 | 120 | 4 | 30 × 25 | 28.32 | 90.90 |
Cutting the temporal tokens from 30 to 8 drops SkeletonMAE from 76.25% to 70.50% and MAMP from 87.00% to 80.90%, while SLiM (batch 768) reaches 90.55%, 90.75% and 90.90% with 8, 15 and 30 tokens. Every row except SLiM-8 (batch 384) is a point of the trade-off plot.
| Method | Tokens | Inference GFLOPs | Single-view module GFLOPs | Reported training resource | |||
|---|---|---|---|---|---|---|---|
| NT × NJ | Enc. | Enc. | Dec. | Target/Aux. | GPU type | Batch size | |
| SkeletonMAE | 30 × 25 | 28.32 | 1.97 | 17.70 | – | 8×A6000 | 64 |
| MAMP | 30 × 25 | 28.32 | 1.97 | 17.70 | – | 4×RTX3090 | 128 |
| S-JEPA | 30 × 25 | 28.32 | 1.97 | 17.70 | 28.32 | 8×A100 | 256 |
| GFP | 30 × 25 | 28.32 | 1.97 | 1.57 | 0.64 | – | – |
| SLiM (Ours) | 8 × 25 | 3.59 | 3.59 | – | 3.59 | 4×A6000 | 768 |
The 7.89× figure concerns the downstream encoder, where every method processes its full unmasked token grid; single-view module GFLOPs are not aggregate pre-training-cost ratios.
Prediction target and robustness Tables 10 and 11
| MFP target | Loss | Target-specific module | X-Sub | X-View | Avg. |
|---|---|---|---|---|---|
| Raw joint coordinates | MSE | 5-block decoder | 84.30 | 89.62 | 86.96 |
| Teacher-projected features | MSE | None | 84.53 | 89.32 | 86.92 |
| Teacher-projected features | MSE | Predictor head | 85.07 | 90.12 | 87.59 |
| Teacher prototype distribution | Cross-entropy | None | 87.66±0.11 | 92.68±0.09 | 90.17 |
Coordinate reconstruction and direct feature regression perform almost identically (86.96 vs. 86.92 average accuracy): removing the decoder or predicting features is not by itself the source of the gain, which the paper attributes to the complete prototype-assignment design.
(a) Declared missingness and noise
| Condition | Severity | X-Sub | X-View |
|---|---|---|---|
| Clean | – | 87.9 | 93.2 |
| Declared missing | 3/25 | 87.86±0.06 | 92.94±0.03 |
| Declared missing | 5/25 | 87.69±0.12 | 92.68±0.09 |
| Declared missing | 8/25 | 87.23±0.01 | 92.14±0.01 |
| Gaussian noise | σ = 0.01 | 87.78±0.04 | 93.18±0.02 |
| Gaussian noise | σ = 0.03 | 87.71±0.05 | 93.01±0.02 |
| Gaussian noise | σ = 0.05 | 87.62±0.02 | 92.84±0.03 |
(b) Missing-joint representation (X-Sub)
| Representation | Missing joints / 25 | ||
|---|---|---|---|
| 3 | 5 | 8 | |
| Declared mask token | 87.86 | 87.69 | 87.23 |
| Parent-joint fill | 81.30±0.29 | 76.41±0.06 | 66.51±0.07 |
| Zero fill | 72.20±0.18 | 62.23±0.21 | 49.63±0.23 |
With eight declared missing joints, accuracy remains 87.23/92.14; presenting parent-joint or zero-filled coordinates as valid input degrades it much more.
Predict features, not coordinates
- 01Decoder-free masked feature prediction. The student predicts the EMA teacher’s features at masked tokens, without a reconstruction decoder.
- 02Semantic Tube Masking. Masks connected body parts over time at a 50–90% ratio, forcing inference from broader body context.
- 03Skeleton-Aware Augmentations. 360° vertical-axis rotation, left–right mirroring and bone-length scaling keep views anatomically valid.
BibTeX
@inproceedings{do2026less,
title={Less is More: Compact-Token Masked Feature Prediction for Skeleton Representation Learning},
author={Do, Jeonghyeok and Chen, Yun and Youk, Geunhyuk and Kim, Munchurl},
booktitle={Advances in Neural Information Processing Systems (NeurIPS)},
year={2026}
}