More research

One for All: Generalist Foundation Model for Cross-Sensor Skeleton Representation Learning

Jeonghyeok Do Yun Chen Munchurl Kim†

Korea Advanced Institute of Science and Technology (KAIST), South Korea

† Corresponding authorehwjdgur0913@kaist.ac.krcyruby@kaist.ac.krmkimee@kaist.ac.kr

arXiv preprint, 2026

Abstract

For learning generalizable motion representations from large-scale unlabeled data, Self-supervised learning (SSL) has become a widely adopted methodology. However, existing approaches are primarily limited by the inherent heterogeneity of skeleton data—characterized by varying joint counts, indexing protocols, and topological structures across different sensors—which typically necessitates training separate, sensor-specific, or even entirely dataset-specific models.

To overcome this, we introduce SOfA (Skeleton One for All), the first generalist foundation model designed to achieve sensor-unified skeleton representation learning across diverse sensors. To accommodate the dimensional gap caused by varying joint counts, we introduce a fixed-size set of learnable Canonical Joint Slots, acting as a universal vessel that seamlessly accommodates arbitrary skeletal topologies. SOfA fills these slots via an attention mechanism that dynamically aggregates skeletal information from sensor-specific inputs. Furthermore, we resolve joint index misalignment between various sensors by introducing a Semantic Joint Embedding derived from a pre-trained text encoder, rather than relying on absolute positional embeddings.

To validate our approach, we standardized ten 3D skeleton datasets for unified training. Extensive experiments demonstrate that SOfA can serve as a truly universal encoder, achieving state-of-the-art (SOTA) performance across a wide range of downstream tasks and sensor types, often outperforming dataset-specific models with a single foundation model.

One encoder for any skeleton sensor

One set of weights for 25-, 20- and unseen 15-joint skeletons from ten datasets, instead of one model per sensor or dataset.

Two panels. (a) Previous methods: skeleton sequence X1 from sensor 1 and X2 from sensor 2 go to separate encoders E1 and E2, which are not compatible with each other, giving sensor-specific representations of size T by J1 and T by J2. (b) SOfA: sequences from various sensors are used as the condition for a single encoder E that fills learnable Canonical Joint Slots S_c of size 1 by J_c, giving a sensor-unified representation of size T by J_c.
Sensor-specific versus sensor-unified. (a) Previous methods need a separate encoder for each sensor topology; (b) SOfA fills a fixed set of learnable Canonical Joint Slots from any sensor with one shared encoder.

Paper figures

The learned Canonical Joint Slots in t-SNE, the leakage-free NTU evaluation protocol, and normalized samples from the ten datasets.

Ten datasets Table 14
Table 14. The ten heterogeneous 3D skeleton datasets. Grouped by tracking protocol and skeletal topology; the 15-joint protocols are reserved exclusively for unseen evaluation.
Scroll for more columns
DatasetTrackerJointsClassesSamplesEvaluation Setting
NTU-60Kinect v2256056,880Pre-train & Downstream
NTU-120Kinect v225120114,480Pre-train & Downstream
PKU-MMDKinect v225517,096Pre-train & Downstream
ETRI-ActKinect v22555112,620Pre-train & Downstream
ETRI-LivingLabKinect v225558,605Pre-train & Downstream
MSR-Action3DKinect v12020567Pre-train & Downstream
NW-UCLAKinect v120101,494Pre-train & Downstream
UT-KinectKinect v12010199Pre-train & Downstream
SBU-InterCustom Tracker 1158282Unseen (OOD)
FlorenceCustom Tracker 2159215Unseen (OOD)

Quantitative results

67.1% on PKU-MMD with SOfA-L (+6.8 points over the previous best), 92.9% on NW-UCLA, and 91.7% / 97.0% on the unseen 15-joint Florence and SBU-Inter.

Table 2. Comparison with recent SSL methods on NTU-60, NTU-120 and PKU-MMD (25-joint).
Bold bestUnderline second bestTop-1 accuracy (%)
Scroll for more columns
MethodPublicationNTU-60NTU-120PKU-MMD
X-SubX-ViewX-SubX-SetX-Sub
Sensor-Specific Representation
GL-TransformerECCV'2276.383.866.068.7–
CPMECCV'2278.784.968.769.648.3
CMDECCV'2279.886.970.371.543.0
AimCLRAAAI'2274.379.763.463.4–
HYSPICLR'2378.282.661.864.6–
HaLPCVPR'2379.786.871.172.243.5
ActCLRCVPR'2380.986.769.070.5–
RVTCLRICCV'2374.779.1–––
SkeletonMAEICMEW'2374.877.772.573.536.1
MAMPICCV'2384.989.178.679.153.8
PTSLAAAI'2377.381.866.267.749.3
S-JEPAECCV'2485.389.879.679.953.5
IGMECCV'2486.291.280.081.4–
MacDiffECCV'2486.491.079.480.2–
HSPCVPR'2580.788.071.073.248.9
USDRLAAAI'2585.291.776.678.154.4
GFPICCV'2585.992.079.180.356.2
AMRCVPR'2687.492.381.181.960.3
Sensor-Unified Representation
SOfA (Ours)85.792.1*79.280.465.4
SOfA-L (Ours)87.092.6*79.582.067.1

X-View*: leakage-free evaluation on the official NTU-60 X-View test set without the 5,392 sequences that also appear in the pre-training corpus (13,568 remain; Fig. 5).

Linear evaluation with the joint modality alone, no multi-stream ensemble. Baselines are trained separately for each dataset; SOfA uses the same weights for every benchmark.

Ablations Tables 5 and 13
Table 5. Ablation of SOfA components. The contributions of CJS and SJE across heterogeneous topologies (Top-1 accuracy, %).
Scroll for more columns
Modules25-Joint Datasets20-Joint Datasets
CJSSJENTU-60NTU-120PKU-MMDNW-UCLAUT-Kinect
✗✗75.270.750.968.977.0
✓✗84.978.764.772.380.0
✗✓76.471.452.386.791.0
✓✓85.779.265.492.997.0

Removing CJS lowers accuracy on every benchmark (average drop of 8.48%); removing SJE hurts the 20-joint datasets most (NW-UCLA 92.9% → 72.3%).

Table 13. Number of Canonical Joint Slots. Linear evaluation accuracy (%) on 25-joint and 20-joint benchmarks, with inference GFLOPs.
Scroll for more columns
Slots25-Joint Datasets20-Joint DatasetsGFLOPs
CJSNTU-60NTU-120PKU-MMDNW-UCLAUT-KinectInference
076.471.452.386.791.03.60
585.779.265.492.997.04.30
1085.579.064.592.798.05.41
1585.378.963.591.897.06.55

Five slots give the best overall accuracy; 10 or 15 slots cost more GFLOPs and lower most benchmarks.

Efficiency and scaling Tables 8 and 7
Table 8. Computational complexity on PKU-MMD. Number of tokens, training and inference GFLOPs, and Top-1 accuracy (X-Sub).
Bold best
Scroll for more columns
MethodTokensGFLOPsPKU-MMD
T × JTrainInferenceX-Sub
Sensor-Specific
SkeletonMAE30 × 2519.6728.3236.1
MAMP30 × 2519.6728.3253.8
S-JEPA30 × 2547.9928.3253.5
GFP30 × 254.1828.3256.2
Sensor-Unified
SOfA8 × (25+5)8.604.3065.4
SOfA-L8 × (25+5)14.907.4567.1

SOfA needs 4.30 inference GFLOPs, a 6.59× reduction from the 28.32 GFLOPs of the MAE-based baselines.

Table 7. Network scalability on the 25-joint benchmarks. Top-1 accuracy (%) under linear evaluation with the joint (J) modality. SOfA is the base model, SOfA (J = 30) is evaluated on an extended synthetic topology, and SOfA-L is the deeper and wider variant.
Bold best
Scroll for more columns
MethodModalityNTU-60NTU-120PKU-MMD
X-SubX-ViewX-SubX-SetX-Sub
SOfAJ85.792.1*79.280.465.4
SOfA (J = 30)J85.591.9*79.080.165.6
SOfA-LJ87.092.6*79.582.067.1

The deeper and wider SOfA-L (12 layers, D = 384) improves every benchmark; the synthetic 30-joint topology performs on par with the standard 25 joints. X-View*: leakage-free NTU-60 X-View test subset (see Table 2).

Semi-supervised, retrieval and 20-joint benchmarks Tables 9–12
Table 9. Semi-supervised results on NTU-60 using only 1% labeled data.
Bold bestUnderline second best
Scroll for more columns
MethodsNTU-60
X-SubX-View
Sensor-Specific
CPM56.757.5
CMD50.653.0
HaLP46.648.7
HiCo54.454.8
UmURL58.158.3
SkeletonMAE54.454.6
MAMP66.068.7
S-JEPA67.569.1
USDRL57.360.7
GFP71.872.9
Sensor-Unified
SOfA71.674.2*
Table 10. Action retrieval results on NTU-60 (X-Sub, X-View).
Bold bestUnderline second best
Scroll for more columns
MethodsNTU-60
X-SubX-View
Sensor-Specific
LongT-GAN39.148.1
P&C50.776.3
ISC62.582.6
HaLP65.883.6
HiCo68.384.8
MAMP62.070.0
GFP70.987.1
Sensor-Unified
SOfA72.186.8*
Table 11. Fully supervised comparison on MSR-Action3D (20-joint), linear evaluation.
Top-1 accuracy (%)
Scroll for more columns
MethodsMSR-Action3D
Fully-supervised
HON4D82.2
Rahmani et al.82.7
Tran et al.84.5
Un-supervised
SOfA82.2
Table 12. Fully supervised comparison on UT-Kinect (20-joint), linear evaluation.
Top-1 accuracy (%)
Scroll for more columns
MethodsUT-Kinect
Fully-supervised
Xia et al.90.9
Devanne et al.91.5
Wang et al.96.5
Un-supervised
SOfA97.0

With 1% of the labels, SOfA is best on X-View and 0.2 points below GFP on X-Sub; in retrieval it is best on X-Sub (72.1%). On UT-Kinect it surpasses all compared fully supervised methods (97.0% vs. 96.5%). X-View*: leakage-free NTU-60 X-View test subset (see Table 2).

Comparison of learning paradigms Table 1
Table 1. Conceptual comparison of skeleton representation learning paradigms. Standard SSL methods are constrained by fixed topologies, and HSP is limited to intra-scene multi-modal fusion using paired data; SOfA generalizes across independent datasets and arbitrary sensor configurations.
Scroll for more columns
PropertyStandard SSLHSP (CVPR'25)SOfA (Ours)
Cross-Sensor Capability✗Intra-scene (Paired)Universal (Unpaired)
Alignment Strategy✗Skeletal interpolationCanonical Joint Slots + Semantic Joint Embedding
Sensor Flexibility✗Pre-defined 2D-3D pairArbitrary 3D sensor
Dataset DependencyHomogeneousHeterogeneous (Paired)Heterogeneous (Independent)
Representation ScopeDomain-specificMulti-modal fusionGeneralist foundation

Any skeleton in, canonical slots out

SOfA framework. Skeleton sequences sampled from various sensors are patchified and given a Semantic Joint Embedding. The resulting patch tokens Z_p are combined with temporally repeated Canonical Joint Slots Z_c. The teacher network receives the full input; the student network receives a masked input and its ViT encoder feeds a canonical head and a DINO head. The canonical loss L_Cano makes the student fill the slots like the teacher, and the DINO loss aligns the class tokens; the teacher is an EMA of the student. The representation W, made of the class token and the canonical tokens, is used for downstream tasks.
Overview of SOfA. Teacher and student both fill a fixed set of Canonical Joint Slots; the student sees a masked view and matches the unmasked teacher, and only the canonical representation is kept for downstream tasks.
  1. 01Canonical Joint Slots. A fixed set of learnable slots is filled from any sensor’s joint tokens by masked self-attention; padded joints are masked out.
  2. 02Semantic Joint Embedding. Joint names such as “Left Wrist” pass through a pretrained T5 text encoder in place of absolute positional embeddings.
  3. 03Teacher–student pre-training. From a 90%-masked view, the student matches the EMA teacher’s slot (LCano) and class-token (LDINO) outputs.

BibTeX

@article{do2026sofa,
  title={One for All: Generalist Foundation Model for Cross-Sensor Skeleton Representation Learning},
  author={Do, Jeonghyeok and Chen, Yun and Kim, Munchurl},
  journal={arXiv preprint arXiv:2609.07078},
  year={2026}
}
Prior work: SLiM (NeurIPS 2026)
@inproceedings{do2026less,
  title={Less is More: Compact-Token Masked Feature Prediction for Skeleton Representation Learning},
  author={Do, Jeonghyeok and Chen, Yun and Youk, Geunhyuk and Kim, Munchurl},
  booktitle={Advances in Neural Information Processing Systems (NeurIPS)},
  year={2026}
}