More research

Less is More: Compact-Token Masked Feature Prediction for Skeleton Representation Learning

Jeonghyeok Do Yun Chen Geunhyuk Youk Munchurl Kim†

Korea Advanced Institute of Science and Technology (KAIST), South Korea

† Corresponding authorehwjdgur0913@kaist.ac.krcyruby@kaist.ac.krrmsgurkjg@kaist.ac.krmkimee@kaist.ac.kr

NeurIPS 2026

Abstract

Current skeleton representation learning paradigms face distinct limitations: Contrastive Learning (CL) often overlooks fine-grained motion details, while Masked Auto-Encoders (MAE) rely on coordinate-level reconstruction. This reconstruction inherently demands dense token sequences and heavy decoders, wasting pre-training computation on discarded components and forcing downstream inference to process dense token grids. To resolve these bottlenecks, we propose SLiM (Skeleton Less is More), a compact-token framework that unifies masked feature prediction and contrastive learning via a shared encoder. By shifting the objective from raw coordinate reconstruction to decoder-free, teacher-guided feature prediction, SLiM breaks the reliance on dense tokenization and enables effective learning with a highly compact token grid. Crucially, to prevent trivial shortcut learning arising from strong inter-joint dependencies of human, we introduce Semantic Tube Masking together with Skeleton-Aware Augmentations to enforce deep skeletal-temporal reasoning and anatomical consistency. Extensive experiments across multiple downstream protocols demonstrate that SLiM achieves state-of-the-art performance while structurally reducing inference computation by 7.89× compared to dense-token MAE baselines.

Fewer tokens, higher accuracy

A compact 8 × 25 token grid: state-of-the-art accuracy with 7.89× less inference computation than dense-token MAEs.

Line plot of NTU-60 accuracy (averaged over X-Sub and X-View) against downstream inference GFLOPs. SLiM-8, SLiM-15 and SLiM-30 stay between 90.55 and 90.90 percent from 3.59 to 28.32 GFLOPs, above the previous best GFP-30 at 28.32 GFLOPs; MAMP and SkeletonMAE drop sharply when their temporal token budget falls from 30 to 15 and 8.
Accuracy–efficiency trade-off on NTU-60. Mean of X-Sub and X-View vs. inference GFLOPs; the suffix is the temporal token budget NT.
Table 3. Token budget and computational cost comparison on NTU-60 (NT × NJ: encoder token grid after patchification).
Bold best
Scroll for more columns
MethodTokensInference (GFLOPs)Training (GFLOPs)
NT × NJEncoderEncoderDecoder
SkeletonMAE30 × 2528.321.9717.70
MAMP30 × 2528.321.9717.70
S-JEPA30 × 2528.321.9717.70
GFP30 × 2528.321.971.57
SLiM (Ours)8 × 253.593.59–

Paper figures

Feature prediction needs neither dense tokens nor a reconstruction decoder; Semantic Tube Masking and Skeleton-Aware Augmentations prevent shortcut learning.

Two rows. (a) MAE: a 120 by J clip is patchified into 30 by J tokens, 90 percent are masked, an encoder and a heavy decoder restore all tokens by coordinate reconstruction, and inference encodes 30 by J tokens at heavy cost. (b) SLiM: a 64 by J clip becomes 8 by J compact tokens, 50 to 90 percent are corrupted, the encoder predicts teacher features (masked prediction) and aligns a contrastive view, and inference encodes 8 by J tokens, 7.89 times lighter.
MAE vs. SLiM. (a) Dense 30 × J tokens and a decoder that reconstructs coordinates; (b) decoder-free feature prediction on compact 8 × J tokens.

Quantitative results

Best on four of five linear-evaluation protocols, three of four semi-supervised settings and both retrieval protocols, using a single 3D joint stream.

Table 1. Linear evaluation on NTU-60, NTU-120, and PKU-MMD II. We report Top-1 accuracy (%) using the joint modality.
Bold bestUnderline second best
Scroll for more columns
MethodVenueNTU-60NTU-120PKU-MMD II
X-SubX-ViewX-SubX-SetX-Sub
Other skeleton SSL baselines
GL-TransformerECCV'2276.383.866.068.7–
MacDiffECCV'2486.491.079.480.2–
IGMECCV'2486.291.280.081.4–
USDRLAAAI'2585.291.776.678.154.4
HSPCVPR'2580.788.071.073.248.9
Contrastive learning (CL)
CPMECCV'2278.784.968.769.648.3
CMDECCV'2279.886.970.371.543.0
AimCLRAAAI'2274.379.763.463.4–
RVTCLRICCV'2374.779.1–––
PTSLAAAI'2377.381.866.267.749.3
HYSPICLR'2378.282.661.864.6–
HaLPCVPR'2379.786.871.172.243.5
ActCLRCVPR'2380.986.769.070.5–
Masked / predictive modeling
SkeletonMAEICMEW'2374.877.772.573.536.1
MAMPICCV'2384.989.178.679.153.8
S-JEPAECCV'2485.389.879.679.953.5
GFPICCV'2585.992.079.180.356.2
AMRCVPR'2687.492.381.181.960.3
Unified MFP + CL
SLiM (Ours)–87.993.281.283.659.7

Including the concurrent AMR, SLiM is best on four of five protocols; on PKU-MMD II it ranks second, 0.6 points below AMR. All results use the joint modality only, without bone or motion streams.

Ablations Tables 4 and 5
Table 4. Ablation studies for the objective functions on NTU-60. ℒCL represents a standard CL loss without temporal diversity.
Scroll for more columns
ObjectiveNTU-60
ℒMFPℒCLℒGLCLX-SubX-View
✗✓✗73.678.9
✗✗✓75.380.1
✓✗✗85.390.6
✓✓✗86.291.3
✓✗✓87.792.7

ℒMFP alone reaches 85.3% on X-Sub; adding a standard CL loss gives 86.2%, and ℒGLCL instead gives 87.7%. Ablation models are pre-trained for 100 epochs (final model: 150).

Table 5. Ablation studies on NTU-60. (a) Dual-role Semantic Tube Masking (STM) for the global masked-view branch and local contrastive-view branch. (b) Skeleton-Aware Augmentations (SAA), including rotation, mirroring, and scaling. Limited denotes the corresponding conventional masking or augmentation strategy.
Scroll for more columns

(a) Semantic Tube Masking

MaskingNTU-60
Global viewLocal viewX-SubX-View
✗✓75.380.1
✓ (Limited)✓84.290.4
✓✗83.389.4
✓✓ (Limited)85.789.9
✓ (90%)✓ (90%)86.892.1
✓ (0–50%)✓ (0–50%)84.589.6
✓ (50–90%)✓ (50–90%)87.792.7

(b) Skeleton-Aware Augmentations

RotationMirroringScalingX-SubX-View
✗✓✓81.686.9
✓ (Limited)✓✓84.889.9
✓✗✓85.390.6
✓✓ (Limited)✓85.991.0
✓✓✗84.490.0
✓✓✓ (Limited)86.491.7
✓✓✓87.792.7

Replacing STM with conventional joint masking in either role lowers X-Sub by up to 3.5 points (87.7 vs. 84.2); the stochastic 50–90% ratio beats a fixed 90% by 0.9 and 0–50% by 3.2 points. Limited is the conventional counterpart shown in Fig. 4 (a)–(d).

Token budget and cost Tables 7 and 8
Table 7. Token-budget configurations and average accuracy on NTU-60. T denotes the sampled clip length, PT denotes the temporal patch size, and NT = T/PT is the number of temporal tokens. All settings use J = 25 joints and PJ = 1. Accuracy is averaged over the X-Sub and X-View protocols under linear evaluation. Inference GFLOPs are measured for the downstream encoder. SLiM-8 uses the default total batch size of 768 unless otherwise specified.
Scroll for more columns
MethodClip TPatch PTTokens NT × NJInference GFLOPsAvg. Acc. (%)
SkeletonMAE-86488 × 253.5970.50
SkeletonMAE-15120815 × 258.3674.00
SkeletonMAE-30120430 × 2528.3276.25
MAMP-86488 × 253.5980.90
MAMP-15120815 × 258.3685.50
MAMP-30120430 × 2528.3287.00
S-JEPA120430 × 2528.3287.55
GFP120430 × 2528.3288.95
SLiM-8 (Batch 384)6488 × 253.5989.60
SLiM-8 (Batch 768)6488 × 253.5990.55
SLiM-15120815 × 258.3690.75
SLiM-30120430 × 2528.3290.90

Cutting the temporal tokens from 30 to 8 drops SkeletonMAE from 76.25% to 70.50% and MAMP from 87.00% to 80.90%, while SLiM (batch 768) reaches 90.55%, 90.75% and 90.90% with 8, 15 and 30 tokens. Every row except SLiM-8 (batch 384) is a point of the trade-off plot.

Table 8. Detailed cost comparison on NTU-60. NT × NJ denotes the encoder token grid after patchification. Inference GFLOPs are measured for the downstream encoder. Module-level GFLOPs are separated into student encoder, decoder, and target/auxiliary branches for one forward view. The target/auxiliary branch includes target generation networks for prior methods when used, and the EMA teacher forward for SLiM. These values do not aggregate SLiM's multi-view student/teacher forwards and therefore are not a total per-iteration pre-training comparison. Training resources are reported as implementation context and are not intended as a controlled hardware comparison.
Bold best
Scroll for more columns
MethodTokensInference GFLOPsSingle-view module GFLOPsReported training resource
NT × NJEnc.Enc.Dec.Target/Aux.GPU typeBatch size
SkeletonMAE30 × 2528.321.9717.70–8×A600064
MAMP30 × 2528.321.9717.70–4×RTX3090128
S-JEPA30 × 2528.321.9717.7028.328×A100256
GFP30 × 2528.321.971.570.64––
SLiM (Ours)8 × 253.593.59–3.594×A6000768

The 7.89× figure concerns the downstream encoder, where every method processes its full unmasked token grid; single-view module GFLOPs are not aggregate pre-training-cost ratios.

Prediction target and robustness Tables 10 and 11
Table 10. Matched comparison of masked-prediction targets on NTU-60. The uncertainty shown for the proposed variant is the sample standard deviation over three independently trained linear probes on one fixed pre-trained checkpoint; it is not pre-training-run variability.
Bold best
Scroll for more columns
MFP targetLossTarget-specific moduleX-SubX-ViewAvg.
Raw joint coordinatesMSE5-block decoder84.3089.6286.96
Teacher-projected featuresMSENone84.5389.3286.92
Teacher-projected featuresMSEPredictor head85.0790.1287.59
Teacher prototype distributionCross-entropyNone87.66±0.1192.68±0.0990.17

Coordinate reconstruction and direct feature regression perform almost identically (86.96 vs. 86.92 average accuracy): removing the decoder or predicting features is not by itself the source of the gain, which the paper attributes to the complete prototype-assignment design.

Table 11. Test-time robustness on NTU-60. (a) Declared missing joints and localization noise, where σ is normalized by each sequence's mean bone length. (b) Unflagged missing joints on X-Sub. The same selected joints are absent throughout a clip. Each ± value is the sample standard deviation over three corruption draws, not pre-training variability.
Scroll for more columns

(a) Declared missingness and noise

ConditionSeverityX-SubX-View
Clean–87.993.2
Declared missing3/2587.86±0.0692.94±0.03
Declared missing5/2587.69±0.1292.68±0.09
Declared missing8/2587.23±0.0192.14±0.01
Gaussian noiseσ = 0.0187.78±0.0493.18±0.02
Gaussian noiseσ = 0.0387.71±0.0593.01±0.02
Gaussian noiseσ = 0.0587.62±0.0292.84±0.03

(b) Missing-joint representation (X-Sub)

RepresentationMissing joints / 25
358
Declared mask token87.8687.6987.23
Parent-joint fill81.30±0.2976.41±0.0666.51±0.07
Zero fill72.20±0.1862.23±0.2149.63±0.23

With eight declared missing joints, accuracy remains 87.23/92.14; presenting parent-joint or zero-filled coordinates as valid input degrades it much more.

Predict features, not coordinates

SLiM framework. Left: multi-view generation, with a primary interval re-sampled into global view XG1 (64 frames) and local views XL1 to XL3 (32, 16 and 8 frames), and a secondary interval giving global view XG2. Views pass Skeleton-Aware Augmentations, patchify and joint embedding. The teacher network, updated by EMA, encodes the unmasked anchor view; Semantic Tube Masking produces the masked tokens for the student network, whose ViT encoder feeds a masked feature prediction head and a contrastive head. Losses: masked feature prediction against the teacher's patch features and global-local contrastive learning against the teacher's class token. The encoder output W is used for downstream tasks.
Overview of SLiM. Masked feature prediction (MFP) and global–local contrastive learning (GLCL) with one compact-token encoder.
  1. 01Decoder-free masked feature prediction. The student predicts the EMA teacher’s features at masked tokens, without a reconstruction decoder.
  2. 02Semantic Tube Masking. Masks connected body parts over time at a 50–90% ratio, forcing inference from broader body context.
  3. 03Skeleton-Aware Augmentations. 360° vertical-axis rotation, left–right mirroring and bone-length scaling keep views anatomically valid.

BibTeX

@inproceedings{do2026less,
  title={Less is More: Compact-Token Masked Feature Prediction for Skeleton Representation Learning},
  author={Do, Jeonghyeok and Chen, Yun and Youk, Geunhyuk and Kim, Munchurl},
  booktitle={Advances in Neural Information Processing Systems (NeurIPS)},
  year={2026}
}