Attend to the key joints and frames
Self-attention within four partitions, neighboring or distant joints × local or global motion: state-of-the-art at 3.62 GFLOPs.
Paper figures
Skate-Type importance scores per action, the Skate-MSA branches, the partition and reverse operations, trimmed-uniform frame sampling and the joint partitions.
Quantitative results
92.6% on X-Sub60 and 97.0% on X-View60 with the joint modality alone, 98.3% on NW-UCLA; best in every setting of Table 1 except the four-modality ensemble on X-Sub120.
| Methods | Frames | NTU RGB+D (%) | NTU RGB+D 120 (%) | NW-UCLA (%) | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| X-Sub60 | X-View60 | X-Sub120 | X-Set120 | ||||||||||||
| πΌ1 | πΌ2 | πΌ4 | πΌ1 | πΌ2 | πΌ4 | πΌ1 | πΌ2 | πΌ4 | πΌ1 | πΌ2 | πΌ4 | ||||
| RNN | |||||||||||||||
| AGC-LSTM | 100 | 87.5 | 89.2 | β | 93.5 | 95.0 | β | β | β | β | β | β | β | 93.3 | |
| CNN | |||||||||||||||
| TA-CNN | 64 | 88.8 | β | 90.4 | 93.6 | β | 94.8 | 82.4 | β | 85.4 | 84.0 | β | 86.8 | 96.1 | |
| Ske2Grid | 100 | 88.3 | β | β | 95.7 | β | β | 82.7 | β | β | 85.1 | β | β | β | |
| GCN | |||||||||||||||
| SGN | 20 | β | 89.0 | β | β | 94.5 | β | β | 79.2 | β | β | 81.5 | β | β | |
| CTR-GCN | 64 | 89.9 | β | 92.4 | β | β | 96.8 | 84.9 | 88.7 | 88.9 | β | 90.1 | 90.6 | 96.5 | |
| ST-GCN++ | 100 | 89.3 | 91.4 | 92.1 | 95.6 | 96.7 | 97.0 | 83.2 | 87.0 | 87.5 | 85.6 | 87.5 | 89.8 | β | |
| InfoGCN | 64 | β | β | 92.7 | β | β | 96.9 | 85.1 | 88.5 | 89.4 | 86.3 | 89.7 | 90.7 | 96.6 | |
| FR-Head | 64 | 90.3 | 92.3 | 92.8 | 95.3 | 96.4 | 96.8 | 85.5 | β | 89.5 | 87.3 | β | 90.9 | 96.8 | |
| Koopman | 64 | 90.2 | β | 92.9 | 95.2 | β | 96.8 | 85.7 | β | 90.0 | 87.4 | β | 91.3 | 97.0 | |
| LST | 64 | 90.2 | β | 92.9 | 95.6 | β | 97.0 | 85.5 | β | 89.9 | 87.0 | β | 91.1 | 97.2 | |
| HD-GCN | 64 | 90.6 | 92.4 | 93.0 | 95.7 | 96.6 | 97.0 | 85.7 | 89.1 | 89.8 | 87.3 | 90.6 | 91.2 | 96.9 | |
| STC-Net | 64 | β | 92.5 | 93.0 | β | 96.7 | 97.1 | β | 89.3 | 89.9 | β | 90.7 | 91.3 | 97.2 | |
| Transformer | |||||||||||||||
| DSTA-Net | 128 | β | β | 91.5 | β | β | 96.4 | β | β | 86.6 | β | β | 89.0 | β | |
| STST | 128 | β | β | 91.9 | β | β | 96.8 | β | β | β | β | β | β | β | |
| FG-STFormer | 128 | β | β | 92.6 | β | β | 96.7 | β | β | 89.0 | β | β | 90.6 | 97.0 | |
| Hyperformer | 64 | 90.7 | β | 92.9 | 95.1 | β | 96.5 | 86.6 | β | 89.9 | 88.0 | β | 91.3 | 96.7 | |
| SkateFormer | 64 | 92.6 | 93.0 | 93.5 | 97.0 | 97.4 | 97.8 | 87.7 | 89.4 | 89.8 | 89.3 | 91.0 | 91.4 | 98.3 | |
𝔼1: joint modality only; 𝔼2: joint + bone; 𝔼4: joint + bone + joint motion + bone motion, with a separate network for each modality and their outputs ensembled.
| Methods | NTU-Inter (πΌ1, %) | NTU-Inter 120 (πΌ1, %) | Params. (M) | FLOPs (G) | Time (ms) | ||
|---|---|---|---|---|---|---|---|
| X-Sub60 | X-View60 | X-Sub120 | X-Set120 | ||||
| Transformer | |||||||
| IGFormer | 93.6 | 96.5 | 85.4 | 86.5 | β | β | β |
| SkeleTR | 94.9 | 97.7 | 87.8 | 88.3 | 3.82 | 7.30 | β |
| ISTA-Net | β | β | 90.6 | 91.7 | 6.22 | 68.18 | 21.71 |
| SkateFormer | 97.1 | 99.3 | 92.3 | 93.2 | 2.02 | 3.62 | 11.25 |
NTU-Inter and NTU-Inter 120 are the 11 and 26 two-person interaction classes of NTU RGB+D and NTU RGB+D 120.
| Methods | Params. (M) | FLOPs (G) | Time (ms) | NTU RGB+D (πΌ1, %) | NTU RGB+D 120 (πΌ1, %) |
|---|---|---|---|---|---|
| GCN | |||||
| InfoGCN | 1.56 | 3.34 | 12.97 | β | 85.7 |
| FR-Head | 1.45 | 3.60 | 18.49 | 92.8 | 86.4 |
| Koopman | 5.38 | 8.76 | 17.86 | 92.7 | 86.6 |
| LST | 2.10 | 3.60 | 18.85 | 92.9 | 86.3 |
| HD-GCN | 1.66 | 3.44 | 72.81 | 93.2 | 86.5 |
| Transformer | |||||
| DSTA-Net | 3.45 | 16.18 | 13.80 | β | β |
| Hyperformer | 2.71 | 9.64 | 18.07 | 92.9 | 87.3 |
| SkateFormer | 2.03 | 3.62 | 11.46 | 94.8 | 88.5 |
Accuracy: average top-1 of the two benchmarks of each dataset, joint modality. Compiled with the publicly available official code of each method.
| Attention Types in attention layer | Total 8 attention layers | First two attention layers | ||
|---|---|---|---|---|
| FLOPs (M) | Memory (MB) | FLOPs (M) | Memory (MB) | |
| vnjpk + vdjpl + tlocalm + tglobaln | 29.12 | 56.3 | 15.53 | 27.0 |
| Naive self-attention | 801.79Β (Γ27.53) | 1530.0Β (Γ27.18) | 603.65Β (Γ38.87) | 1152.0Β (Γ42.67) |
The 48× reduction stated in the paper is a theoretical figure for the first two SkateFormer blocks, before downsampling, where the effect is largest; measured on the attention layers alone, Skate-MSA needs 38.87× fewer FLOPs there.
| Methods | Frames | NTU RGB+D (%) | |||||||
|---|---|---|---|---|---|---|---|---|---|
| X-Sub60 | X-View60 | ||||||||
| J | B | JM | BM | J | B | JM | BM | ||
| ST-GCN (*) | 100 | 87.8 | 88.6 | 85.8 | 86.2 | 95.5 | 95.0 | 93.7 | 92.8 |
| CTR-GCN | 64 | 89.9 | 90.6 | 88.1 | 87.9 | β | β | β | β |
| CTR-GCN (*) | 100 | 89.6 | 90.0 | 88.0 | 87.5 | 95.6 | 95.4 | 94.4 | 93.6 |
| ST-GCN++ | 100 | 89.3 | 90.1 | 87.5 | 87.3 | 95.6 | 95.5 | 94.3 | 93.8 |
| InfoGCN | 64 | 89.8 | 90.6 | 88.9 | 88.6 | 95.2 | 95.5 | 94.2 | 93.6 |
| FR-Head | 64 | 90.3 | 91.1 | 88.7 | 87.6 | 95.3 | 95.0 | 93.6 | 92.6 |
| LST | 64 | 90.2 | 91.2 | 88.0 | 87.8 | 95.6 | 95.5 | 93.7 | 93.2 |
| HD-GCN | 64 | 90.6 | 90.9 | β | β | 95.7 | 95.1 | β | β |
| SkateFormer | 64 | 92.6 | 92.1 | 89.8 | 89.0 | 97.0 | 96.5 | 95.8 | 94.7 |
J: joint, B: bone, JM: joint motion, BM: bone motion. (*): accuracy as reported in ST-GCN++, since the official code and paper give no modality-specific numbers.
| Methods | Frames | NTU RGB+D 120 (%) | |||||||
|---|---|---|---|---|---|---|---|---|---|
| X-Sub120 | X-Set120 | ||||||||
| J | B | JM | BM | J | B | JM | BM | ||
| ST-GCN (*) | 100 | 82.1 | 83.7 | 80.3 | 80.6 | 84.5 | 85.8 | 82.7 | 83.0 |
| CTR-GCN | 64 | 84.9 | 85.7 | 81.4 | 81.2 | β | 87.5 | β | β |
| CTR-GCN (*) | 100 | 84.0 | 85.9 | 81.1 | 82.2 | 85.9 | 87.4 | 84.1 | 83.9 |
| ST-GCN++ | 100 | 83.2 | 85.6 | 80.4 | 81.5 | 85.6 | 87.5 | 84.3 | 83.0 |
| InfoGCN | 64 | 85.1 | 87.3 | 82.1 | 82.5 | 86.3 | 88.5 | 84.4 | 84.8 |
| FR-Head | 64 | 85.5 | 86.8 | 81.9 | 82.0 | 87.3 | 88.1 | 84.0 | 83.9 |
| LST | 64 | 85.5 | 87.5 | 82.3 | 82.4 | 87.0 | 88.7 | 83.9 | 84.4 |
| HD-GCN | 64 | 85.7 | 86.7 | β | β | 87.3 | 88.4 | β | β |
| SkateFormer | 64 | 87.7 | 88.2 | 83.1 | 82.3 | 89.3 | 89.8 | 85.3 | 84.1 |
Ablations Tables 4, 5 and 6
| Attention Types | NTU RGB+D (%) | Params. (M) | FLOPs (G) | Time (ms) | Memory (MB) | |
|---|---|---|---|---|---|---|
| X-Sub60 | X-View60 | |||||
| Baseline (no attention) | 90.7 | 95.7 | 2.03 | 3.59 | 10.68 | 80.6 |
| + Naive self-attention* | 90.6Β (β0.1) | 95.0Β (β0.7) | 2.03 | 4.39 | 17.55 | 1612.1 |
| + vnjpk + vdjpl | 91.8Β (β1.1) | 96.4Β (β0.7) | 2.03 | 3.59 | 11.10 | 86.9 |
| + tlocalm + tglobaln | 91.9Β (β1.2) | 96.6Β (β0.9) | 2.03 | 3.59 | 11.09 | 85.5 |
| + vnjpk + vdjpl + tlocalm + tglobaln | 92.6Β (β1.9) | 97.0Β (β1.3) | 2.03 | 3.62 | 11.46 | 137.6 |
Each relation type helps on its own, and together they give the best accuracy (+1.9 / +1.3 points over the baseline). With Skate-MSA the model needs 137.6 MB, against 80.6 MB without attention and 1612.1 MB with naive self-attention over all joints and frames (run with batch size 8).
| Embedding Methods | NTU RGB+D (%) | ||
|---|---|---|---|
| Skeletal | Temporal | X-Sub60 | X-View60 |
| β | β | 91.9Β (β0.7) | 96.6Β (β0.4) |
| Learnable | 91.8Β (β0.8) | 96.5Β (β0.5) | |
| Fixed (TE) | 91.6Β (β1.0) | 96.6Β (β0.4) | |
| Learnable (SE) | β | 92.1Β (β0.5) | 96.5Β (β0.5) |
| Learnable | 91.4Β (β1.2) | 96.2Β (β0.8) | |
| Fixed (TE) | 92.6 | 97.0 | |
| Intra-instance | Inter-instance | NTU RGB+D (%) | ||
|---|---|---|---|---|
| Temporal | Skeletal | X-Sub60 | X-View60 | |
| Trimmed | 89.8 | 94.6 | ||
| Trimmed | β | 90.3Β (β0.5) | 94.3Β (β0.3) | |
| Trimmed | β | 92.2Β (β2.4) | 96.6Β (β2.0) | |
| Fixed | β | β | 91.4Β (β1.6) | 96.4Β (β1.8) |
| Uniform | β | β | 91.6Β (β1.8) | 96.8Β (β2.2) |
| Trimmed | β | β | 92.6Β (β2.8) | 97.0Β (β2.4) |
Learnable skeletal features with fixed temporal index features (the Skate-Embedding) work best. Trimmed-uniform random sampling surpasses fixed-stride and uniform random sampling by 1.2 (0.6) and 1.0 (0.2) points on X-Sub60 (X-View60).
Per-class accuracy Table 8
| Action Label | Baseline (Rank) | +βvnjpkβ+βvdjpl | +βtlocalmβ+βtglobaln | +βvnjpkβ+βvdjplβ+βtlocalmβ+βtglobaln |
|---|---|---|---|---|
| drink water | 85.77Β (47) | 85.77Β (β0.00) | 85.77Β (β0.00) | 88.32Β (β2.55) |
| eat meal/snack | 76.00Β (56) | 77.82Β (β1.82) | 77.45Β (β1.45) | 78.18Β (β2.18) |
| brushing teeth | 90.11Β (42) | 85.35Β (β4.76) | 91.21Β (β1.10) | 89.01Β (β1.10) |
| brushing hair | 89.01Β (45) | 91.21Β (β2.20) | 93.77Β (β4.76) | 93.41Β (β4.40) |
| drop | 92.00Β (39) | 94.18Β (β2.18) | 92.36Β (β0.36) | 92.73Β (β0.73) |
| pickup | 97.09Β (18) | 97.82Β (β0.73) | 97.09Β (β0.00) | 98.55Β (β1.45) |
| throw | 92.73Β (36) | 93.45Β (β0.73) | 92.00Β (β0.73) | 93.82Β (β1.09) |
| sitting down | 98.17Β (13) | 98.53Β (β0.37) | 99.27Β (β1.10) | 98.53Β (β0.37) |
| standing up (from sitting position) | 98.90Β (7) | 98.53Β (β0.37) | 99.63Β (β0.73) | 98.90Β (β0.00) |
| clapping | 83.88Β (51) | 81.68Β (β2.20) | 86.08Β (β2.20) | 85.35Β (β1.47) |
| reading | 58.61Β (60) | 60.81Β (β2.20) | 61.54Β (β2.93) | 63.37Β (β4.76) |
| writing | 68.38Β (57) | 68.01Β (β0.37) | 66.91Β (β1.47) | 71.32Β (β2.94) |
| tear up paper | 95.94Β (21) | 96.68Β (β0.74) | 94.46Β (β1.48) | 95.94Β (β0.00) |
| wear jacket | 98.18Β (12) | 98.18Β (β0.00) | 98.55Β (β0.36) | 97.82Β (β0.36) |
| take off jacket | 98.19Β (10) | 98.91Β (β0.72) | 99.64Β (β1.45) | 99.64Β (β1.45) |
| wear a shoe | 60.81Β (59) | 86.81Β (β26.01) | 84.62Β (β23.81) | 87.55Β (β26.74) |
| take off a shoe | 78.83Β (53) | 82.48Β (β3.65) | 78.47Β (β0.36) | 83.94Β (β5.11) |
| wear on glasses | 93.04Β (34) | 93.41Β (β0.37) | 93.77Β (β0.73) | 93.41Β (β0.37) |
| take off glasses | 95.62Β (23) | 95.62Β (β0.00) | 95.26Β (β0.36) | 94.89Β (β0.73) |
| put on a hat/cap | 96.69Β (19) | 97.43Β (β0.74) | 98.16Β (β1.47) | 97.79Β (β1.10) |
| take off a hat/cap | 98.90Β (7) | 98.53Β (β0.37) | 98.90Β (β0.00) | 98.90Β (β0.00) |
| cheer up | 93.80Β (29) | 92.34Β (β1.46) | 93.80Β (β0.00) | 94.89Β (β1.09) |
| hand waving | 93.80Β (29) | 94.16Β (β0.36) | 93.80Β (β0.00) | 93.07Β (β0.73) |
| kicking something | 94.20Β (27) | 94.93Β (β0.72) | 97.10Β (β2.90) | 97.83Β (β3.62) |
| reach into pocket | 84.67Β (49) | 84.31Β (β0.36) | 86.13Β (β1.46) | 86.86Β (β2.19) |
| hopping (one foot jumping) | 98.91Β (6) | 98.91Β (β0.00) | 98.91Β (β0.00) | 98.91Β (β0.00) |
| jump up | 100.00Β (1) | 100.00Β (β0.00) | 100.00Β (β0.00) | 100.00Β (β0.00) |
| make a phone call/answer phone | 90.18Β (41) | 89.09Β (β1.09) | 90.55Β (β0.36) | 92.00Β (β1.82) |
| playing with phone/tablet | 76.36Β (55) | 76.36Β (β0.00) | 77.09Β (β0.73) | 74.91Β (β1.45) |
| typing on a keyboard | 64.36Β (58) | 72.73Β (β8.36) | 74.55Β (β10.18) | 77.82Β (β13.45) |
| pointing to something with finger | 79.71Β (52) | 88.04Β (β8.33) | 81.88Β (β2.17) | 84.42Β (β4.71) |
| taking a selfie | 92.39Β (38) | 93.84Β (β1.45) | 94.57Β (β2.17) | 93.12Β (β0.72) |
| check time (from watch) | 93.12Β (33) | 91.67Β (β1.45) | 92.75Β (β0.36) | 92.03Β (β1.09) |
| rub two hands together | 89.13Β (43) | 89.49Β (β0.36) | 91.30Β (β2.17) | 92.03Β (β2.90) |
| nod head/bow | 97.46Β (14) | 98.55Β (β1.09) | 98.55Β (β1.09) | 99.28Β (β1.81) |
| shake head | 94.55Β (26) | 95.27Β (β0.73) | 96.73Β (β2.18) | 96.00Β (β1.45) |
| wipe face | 85.87Β (46) | 87.32Β (β1.45) | 87.68Β (β1.81) | 90.94Β (β5.07) |
| salute | 93.48Β (31) | 93.48Β (β0.00) | 93.84Β (β0.36) | 94.93Β (β1.45) |
| put the palms together | 97.46Β (14) | 97.46Β (β0.00) | 96.74Β (β0.72) | 96.38Β (β1.09) |
| cross hands in front (say stop) | 96.01Β (20) | 97.10Β (β1.09) | 96.38Β (β0.36) | 97.46Β (β1.45) |
| sneeze/cough | 78.62Β (54) | 82.97Β (β4.35) | 83.33Β (β4.71) | 85.51Β (β6.88) |
| staggering | 99.28Β (4) | 99.64Β (β0.36) | 100.00Β (β0.72) | 99.28Β (β0.00) |
| falling | 99.64Β (2) | 100.00Β (β0.36) | 99.64Β (β0.00) | 100.00Β (β0.36) |
| touch head (headache) | 84.06Β (50) | 89.13Β (β5.07) | 84.06Β (β0.00) | 88.41Β (β4.35) |
| touch chest (stomachache/heart pain) | 93.84Β (28) | 94.57Β (β0.72) | 93.48Β (β0.36) | 96.01Β (β2.17) |
| touch back (backache) | 95.29Β (24) | 94.57Β (β0.72) | 95.29Β (β0.00) | 95.65Β (β0.36) |
| touch neck (neckache) | 90.22Β (40) | 90.58Β (β0.36) | 92.39Β (β2.17) | 93.12Β (β2.90) |
| nausea or vomiting condition | 85.45Β (48) | 83.64Β (β1.82) | 84.00Β (β1.45) | 85.45Β (β0.00) |
| use a fan (with hand or paper)/feeling warm | 89.09Β (44) | 90.91Β (β1.82) | 89.45Β (β0.36) | 94.18Β (β5.09) |
| punching/slapping other person | 92.70Β (37) | 93.07Β (β0.36) | 92.70Β (β0.00) | 94.16Β (β1.46) |
| kicking other person | 95.65Β (22) | 96.38Β (β0.72) | 95.65Β (β0.00) | 96.38Β (β0.72) |
| pushing other person | 98.55Β (9) | 97.83Β (β0.72) | 97.46Β (β1.09) | 98.19Β (β0.36) |
| pat on back of other person | 93.48Β (31) | 95.29Β (β1.81) | 94.93Β (β1.45) | 95.65Β (β2.17) |
| point finger at the other person | 92.75Β (35) | 94.20Β (β1.45) | 94.57Β (β1.81) | 94.20Β (β1.45) |
| hugging other person | 99.27Β (5) | 99.27Β (β0.00) | 99.64Β (β0.36) | 99.64Β (β0.36) |
| giving something to other person | 95.29Β (24) | 96.74Β (β1.45) | 95.65Β (β0.36) | 95.65Β (β0.36) |
| touch other person's pocket | 97.45Β (17) | 97.45Β (β0.00) | 98.55Β (β1.09) | 96.00Β (β1.45) |
| handshaking | 97.46Β (14) | 97.10Β (β0.36) | 97.46Β (β0.00) | 97.46Β (β0.00) |
| walking towards each other | 99.63Β (3) | 100.00Β (β0.37) | 100.00Β (β0.37) | 100.00Β (β0.37) |
| walking apart from each other | 98.19Β (10) | 97.10Β (β1.09) | 96.74Β (β1.45) | 97.83Β (β0.36) |
| average | 90.65 | 91.79Β (β1.14) | 91.88Β (β1.23) | 92.62Β (β1.98) |
Compared with the baseline without attention, the full model improves 42 of 60 classes, keeps 8 and slightly lowers 10; ‘wear a shoe’ rises from 60.81% to 87.55%. Most failure cases involve fine finger motions such as reading, writing and typing on a keyboard.
Partition-based transformers compared Table 7
| Methods | Partition Types | Tokenization for Attention | Attention Types |
|---|---|---|---|
| Human Action Recognition | |||
| DSTA-Net | N/A (Reshaping) | No | S-Attn, T-Attn (Sequential) |
| ST-TR | N/A (Reshaping) | No | S-Attn, T-Attn (Parallel) |
| IIP-Transformer | S-Type | Yes | S-Attn, T-Attn (Sequential) |
| STST | N/A (Reshaping) | No | S-Attn, T-Attn (Parallel) |
| FG-STFormer | S-Type | Yes | S-Attn, T-Attn (Sequential) |
| Hyperformer | S-Type | Yes | S-Attn, T-Conv (Sequential) |
| Human Interaction Recognition | |||
| IGFormer | S-Type, T-Type | Yes | Joint-group-level ST-Attn |
| SkeleTR | N/A (Pooling) | Yes | Joint-group-level ST-Attn |
| ISTA-Net | S-Type, T-Type | Yes | Joint-group-level ST-Attn |
| Both | |||
| SkateFormer | 4 Skate-Types | No | Joint-element-level ST-Attn |
Partition, attend, reverse
- 01Four Skate-Types. Joints form neighboring partitions (limbs, torso) and distant partitions (same-position joints across them); frames form local segments of consecutive frames and N-strided global axes. Their four combinations are Skate-Type-1 to -4.
- 02Skate-MSA. Half of the channels go to self-attention, split into four parts that each attend within one Skate-Type partition. In the first blocks this is about 48× less computation than naive self-attention (38.87× measured).
- 03Skate-Embedding. The outer product of learnable skeletal features and fixed sinusoidal features of the sampled frame indices, trained with trimmed-uniform frame sampling and Bone Length AdaIN, which exchanges bone lengths between subjects.
BibTeX
@inproceedings{do2024skateformer,
title={Skateformer: skeletal-temporal transformer for human action recognition},
author={Do, Jeonghyeok and Kim, Munchurl},
booktitle={European Conference on Computer Vision},
pages={401--420},
year={2024},
organization={Springer}
}