More research

Skeletal-Temporal Transformer for Human Action Recognition

Jeonghyeok Do Munchurl Kim†

Korea Advanced Institute of Science and Technology (KAIST), South Korea

† Corresponding authorehwjdgur0913@kaist.ac.krmkimee@kaist.ac.kr

ECCV 2024

Abstract

Skeleton-based action recognition, which classifies human actions based on the coordinates of joints and their connectivity within skeleton data, is widely utilized in various scenarios. While Graph Convolutional Networks (GCNs) have been proposed for skeleton data represented as graphs, they suffer from limited receptive fields constrained by joint connectivity. To address this limitation, recent advancements have introduced transformer-based methods. However, capturing correlations between all joints in all frames requires substantial memory resources. To alleviate this, we propose a novel approach called Skeletal-Temporal Transformer (SkateFormer) that partitions joints and frames based on different types of skeletal-temporal relation (Skate-Type) and performs skeletal-temporal self-attention (Skate-MSA) within each partition. We categorize the key skeletal-temporal relations for action recognition into a total of four distinct types. These types combine (i) two skeletal relation types based on physically neighboring and distant joints, and (ii) two temporal relation types based on neighboring and distant frames. Through this partition-specific attention strategy, our SkateFormer can selectively focus on key joints and frames crucial for action recognition in an action-adaptive manner with efficient computation. Extensive experiments on various benchmark datasets validate that our SkateFormer outperforms recent state-of-the-art methods.

Attend to the key joints and frames

Self-attention within four partitions, neighboring or distant joints × local or global motion: state-of-the-art at 3.62 GFLOPs.

Left: a 3D skeleton whose joints are grouped by dashed outlines into neighboring joint partitions (arms, legs, torso). Middle, skeletal relation types: for 'make OK sign' the key joints are physically neighboring joints, the fingers of one hand; for 'clap' they are physically distant joints, both palms, linked across partitions by red attention arrows over four frames. Right, temporal relation types: for 'brush teeth' the key frames are neighboring frames (local motion); for 'sit down' they are distant frames, the first and the last (global motion). Each example shows the RGB frame next to its skeleton sequence.
Partition-specific attention. Key joints can be physically neighboring (fingers of one hand in ‘make OK sign’) or distant (both palms in ‘clap’); key frames can be neighboring (local motion in ‘brush teeth’) or distant (global motion in ‘sit down’). Left: the neighboring joint partitions.

Paper figures

Skate-Type importance scores per action, the Skate-MSA branches, the partition and reverse operations, trimmed-uniform frame sampling and the joint partitions.

Quantitative results

92.6% on X-Sub60 and 97.0% on X-View60 with the joint modality alone, 98.3% on NW-UCLA; best in every setting of Table 1 except the four-modality ensemble on X-Sub120.

Table 1. Top-1 accuracy of different skeleton-based action recognition methods on the NTU RGB+D, NTU RGB+D 120 and NW-UCLA datasets. The methods of RNN, CNN, GCN, and Transformer types are evaluated based on the number of input frames and the use of ensemble strategies. The best performances are highlighted in bold.
Bold best
Scroll for more columns
MethodsFramesNTU RGB+D (%)NTU RGB+D 120 (%)NW-UCLA
(%)
X-Sub60X-View60X-Sub120X-Set120
𝔼1𝔼2𝔼4𝔼1𝔼2𝔼4𝔼1𝔼2𝔼4𝔼1𝔼2𝔼4
RNN
AGC-LSTM10087.589.2–93.595.0–––––––93.3
CNN
TA-CNN6488.8–90.493.6–94.882.4–85.484.0–86.896.1
Ske2Grid10088.3––95.7––82.7––85.1–––
GCN
SGN20–89.0––94.5––79.2––81.5––
CTR-GCN6489.9–92.4––96.884.988.788.9–90.190.696.5
ST-GCN++10089.391.492.195.696.797.083.287.087.585.687.589.8–
InfoGCN64––92.7––96.985.188.589.486.389.790.796.6
FR-Head6490.392.392.895.396.496.885.5–89.587.3–90.996.8
Koopman6490.2–92.995.2–96.885.7–90.087.4–91.397.0
LST6490.2–92.995.6–97.085.5–89.987.0–91.197.2
HD-GCN6490.692.493.095.796.697.085.789.189.887.390.691.296.9
STC-Net64–92.593.0–96.797.1–89.389.9–90.791.397.2
Transformer
DSTA-Net128––91.5––96.4––86.6––89.0–
STST128––91.9––96.8–––––––
FG-STFormer128––92.6––96.7––89.0––90.697.0
Hyperformer6490.7–92.995.1–96.586.6–89.988.0–91.396.7
SkateFormer6492.693.093.597.097.497.887.789.489.889.391.091.498.3

𝔼1: joint modality only; 𝔼2: joint + bone; 𝔼4: joint + bone + joint motion + bone motion, with a separate network for each modality and their outputs ensembled.

Ablations Tables 4, 5 and 6
Table 4. Ablation study on the influence of skeletal and temporal relation types (SkateTypes) of Skate-MSA in our SkateFormer. (*: The batch size is set to 8 due to limited GPU memory.)
Bold best
Scroll for more columns
Attention TypesNTU RGB+D (%)Params.
(M)
FLOPs
(G)
Time
(ms)
Memory
(MB)
X-Sub60X-View60
Baseline (no attention)90.795.72.033.5910.6880.6
+ Naive self-attention*90.6Β (↓0.1)95.0Β (↓0.7)2.034.3917.551612.1
+ vnjpk + vdjpl91.8Β (↑1.1)96.4Β (↑0.7)2.033.5911.1086.9
+ tlocalm + tglobaln91.9Β (↑1.2)96.6Β (↑0.9)2.033.5911.0985.5
+ vnjpk + vdjpl + tlocalm + tglobaln92.6Β (↑1.9)97.0Β (↑1.3)2.033.6211.46137.6

Each relation type helps on its own, and together they give the best accuracy (+1.9 / +1.3 points over the baseline). With Skate-MSA the model needs 137.6 MB, against 80.6 MB without attention and 1612.1 MB with naive self-attention over all joints and frames (run with batch size 8).

Table 5. Ablation study results on the Skate-Embedding to assess skeletal and temporal embedding impacts.
Bold best
Scroll for more columns
Embedding MethodsNTU RGB+D (%)
SkeletalTemporalX-Sub60X-View60
βœ—βœ—91.9Β (↓0.7)96.6Β (↓0.4)
Learnable91.8Β (↓0.8)96.5Β (↓0.5)
Fixed (TE)91.6Β (↓1.0)96.6Β (↓0.4)
Learnable (SE)βœ—92.1Β (↓0.5)96.5Β (↓0.5)
Learnable91.4Β (↓1.2)96.2Β (↓0.8)
Fixed (TE)92.697.0
Table 6. Ablation study results on the effect of intra-instance and inter-instance augmentations.
Bold best
Scroll for more columns
Intra-instanceInter-instanceNTU RGB+D (%)
TemporalSkeletalX-Sub60X-View60
Trimmed89.894.6
Trimmedβœ“90.3Β (↑0.5)94.3Β (↓0.3)
Trimmedβœ“92.2Β (↑2.4)96.6Β (↑2.0)
Fixedβœ“βœ“91.4Β (↑1.6)96.4Β (↑1.8)
Uniformβœ“βœ“91.6Β (↑1.8)96.8Β (↑2.2)
Trimmedβœ“βœ“92.6Β (↑2.8)97.0Β (↑2.4)

Learnable skeletal features with fixed temporal index features (the Skate-Embedding) work best. Trimmed-uniform random sampling surpasses fixed-stride and uniform random sampling by 1.2 (0.6) and 1.0 (0.2) points on X-Sub60 (X-View60).

Per-class accuracy Table 8
Table 8. Top-1 accuracy based on action labels in NTU RGB+D X-Sub60 evaluation.
Scroll for more columns
Action LabelBaseline (Rank)+ vnjpk + vdjpl+ tlocalm + tglobaln+ vnjpk + vdjpl + tlocalm + tglobaln
drink water85.77Β (47)85.77Β (–0.00)85.77Β (–0.00)88.32Β (↑2.55)
eat meal/snack76.00Β (56)77.82Β (↑1.82)77.45Β (↑1.45)78.18Β (↑2.18)
brushing teeth90.11Β (42)85.35Β (↓4.76)91.21Β (↑1.10)89.01Β (↓1.10)
brushing hair89.01Β (45)91.21Β (↑2.20)93.77Β (↑4.76)93.41Β (↑4.40)
drop92.00Β (39)94.18Β (↑2.18)92.36Β (↑0.36)92.73Β (↑0.73)
pickup97.09Β (18)97.82Β (↑0.73)97.09Β (–0.00)98.55Β (↑1.45)
throw92.73Β (36)93.45Β (↑0.73)92.00Β (↓0.73)93.82Β (↑1.09)
sitting down98.17Β (13)98.53Β (↑0.37)99.27Β (↑1.10)98.53Β (↑0.37)
standing up (from sitting position)98.90Β (7)98.53Β (↓0.37)99.63Β (↑0.73)98.90Β (–0.00)
clapping83.88Β (51)81.68Β (↓2.20)86.08Β (↑2.20)85.35Β (↑1.47)
reading58.61Β (60)60.81Β (↑2.20)61.54Β (↑2.93)63.37Β (↑4.76)
writing68.38Β (57)68.01Β (↓0.37)66.91Β (↓1.47)71.32Β (↑2.94)
tear up paper95.94Β (21)96.68Β (↑0.74)94.46Β (↓1.48)95.94Β (–0.00)
wear jacket98.18Β (12)98.18Β (–0.00)98.55Β (↑0.36)97.82Β (↓0.36)
take off jacket98.19Β (10)98.91Β (↑0.72)99.64Β (↑1.45)99.64Β (↑1.45)
wear a shoe60.81Β (59)86.81Β (↑26.01)84.62Β (↑23.81)87.55Β (↑26.74)
take off a shoe78.83Β (53)82.48Β (↑3.65)78.47Β (↑0.36)83.94Β (↑5.11)
wear on glasses93.04Β (34)93.41Β (↑0.37)93.77Β (↑0.73)93.41Β (↑0.37)
take off glasses95.62Β (23)95.62Β (–0.00)95.26Β (↓0.36)94.89Β (↓0.73)
put on a hat/cap96.69Β (19)97.43Β (↑0.74)98.16Β (↑1.47)97.79Β (↑1.10)
take off a hat/cap98.90Β (7)98.53Β (↓0.37)98.90Β (–0.00)98.90Β (–0.00)
cheer up93.80Β (29)92.34Β (↓1.46)93.80Β (–0.00)94.89Β (↑1.09)
hand waving93.80Β (29)94.16Β (↑0.36)93.80Β (–0.00)93.07Β (↓0.73)
kicking something94.20Β (27)94.93Β (↑0.72)97.10Β (↑2.90)97.83Β (↑3.62)
reach into pocket84.67Β (49)84.31Β (↓0.36)86.13Β (↑1.46)86.86Β (↑2.19)
hopping (one foot jumping)98.91Β (6)98.91Β (–0.00)98.91Β (–0.00)98.91Β (–0.00)
jump up100.00Β (1)100.00Β (–0.00)100.00Β (–0.00)100.00Β (–0.00)
make a phone call/answer phone90.18Β (41)89.09Β (↓1.09)90.55Β (↑0.36)92.00Β (↑1.82)
playing with phone/tablet76.36Β (55)76.36Β (–0.00)77.09Β (↑0.73)74.91Β (↓1.45)
typing on a keyboard64.36Β (58)72.73Β (↑8.36)74.55Β (↑10.18)77.82Β (↑13.45)
pointing to something with finger79.71Β (52)88.04Β (↑8.33)81.88Β (↑2.17)84.42Β (↑4.71)
taking a selfie92.39Β (38)93.84Β (↑1.45)94.57Β (↑2.17)93.12Β (↑0.72)
check time (from watch)93.12Β (33)91.67Β (↓1.45)92.75Β (↓0.36)92.03Β (↓1.09)
rub two hands together89.13Β (43)89.49Β (↑0.36)91.30Β (↑2.17)92.03Β (↑2.90)
nod head/bow97.46Β (14)98.55Β (↑1.09)98.55Β (↑1.09)99.28Β (↑1.81)
shake head94.55Β (26)95.27Β (↑0.73)96.73Β (↑2.18)96.00Β (↑1.45)
wipe face85.87Β (46)87.32Β (↑1.45)87.68Β (↑1.81)90.94Β (↑5.07)
salute93.48Β (31)93.48Β (–0.00)93.84Β (↑0.36)94.93Β (↑1.45)
put the palms together97.46Β (14)97.46Β (–0.00)96.74Β (↓0.72)96.38Β (↓1.09)
cross hands in front (say stop)96.01Β (20)97.10Β (↑1.09)96.38Β (↑0.36)97.46Β (↑1.45)
sneeze/cough78.62Β (54)82.97Β (↑4.35)83.33Β (↑4.71)85.51Β (↑6.88)
staggering99.28Β (4)99.64Β (↑0.36)100.00Β (↑0.72)99.28Β (–0.00)
falling99.64Β (2)100.00Β (↑0.36)99.64Β (–0.00)100.00Β (↑0.36)
touch head (headache)84.06Β (50)89.13Β (↑5.07)84.06Β (–0.00)88.41Β (↑4.35)
touch chest (stomachache/heart pain)93.84Β (28)94.57Β (↑0.72)93.48Β (↓0.36)96.01Β (↑2.17)
touch back (backache)95.29Β (24)94.57Β (↓0.72)95.29Β (–0.00)95.65Β (↑0.36)
touch neck (neckache)90.22Β (40)90.58Β (↑0.36)92.39Β (↑2.17)93.12Β (↑2.90)
nausea or vomiting condition85.45Β (48)83.64Β (↓1.82)84.00Β (↓1.45)85.45Β (–0.00)
use a fan (with hand or paper)/feeling warm89.09Β (44)90.91Β (↑1.82)89.45Β (↑0.36)94.18Β (↑5.09)
punching/slapping other person92.70Β (37)93.07Β (↑0.36)92.70Β (–0.00)94.16Β (↑1.46)
kicking other person95.65Β (22)96.38Β (↑0.72)95.65Β (–0.00)96.38Β (↑0.72)
pushing other person98.55Β (9)97.83Β (↓0.72)97.46Β (↓1.09)98.19Β (↓0.36)
pat on back of other person93.48Β (31)95.29Β (↑1.81)94.93Β (↑1.45)95.65Β (↑2.17)
point finger at the other person92.75Β (35)94.20Β (↑1.45)94.57Β (↑1.81)94.20Β (↑1.45)
hugging other person99.27Β (5)99.27Β (–0.00)99.64Β (↑0.36)99.64Β (↑0.36)
giving something to other person95.29Β (24)96.74Β (↑1.45)95.65Β (↑0.36)95.65Β (↑0.36)
touch other person's pocket97.45Β (17)97.45Β (–0.00)98.55Β (↑1.09)96.00Β (↓1.45)
handshaking97.46Β (14)97.10Β (↓0.36)97.46Β (–0.00)97.46Β (–0.00)
walking towards each other99.63Β (3)100.00Β (↑0.37)100.00Β (↑0.37)100.00Β (↑0.37)
walking apart from each other98.19Β (10)97.10Β (↓1.09)96.74Β (↓1.45)97.83Β (↓0.36)
average90.6591.79Β (↑1.14)91.88Β (↑1.23)92.62Β (↑1.98)

Compared with the baseline without attention, the full model improves 42 of 60 classes, keeps 8 and slightly lowers 10; ‘wear a shoe’ rises from 60.81% to 87.55%. Most failure cases involve fine finger motions such as reading, writing and typing on a keyboard.

Partition-based transformers compared Table 7
Table 7. Comparison with ours SkateFormer with existing transformer-based methods. S-Attn, T-Attn, T-Conv, and ST-Attn indicate β€˜skeletal attention’, β€˜temporal attention’, β€˜temporal convolution’ and β€˜skeletal-temporal attention’, respectively.
Scroll for more columns
MethodsPartition TypesTokenization
for Attention
Attention Types
Human Action Recognition
DSTA-NetN/A (Reshaping)NoS-Attn, T-Attn (Sequential)
ST-TRN/A (Reshaping)NoS-Attn, T-Attn (Parallel)
IIP-TransformerS-TypeYesS-Attn, T-Attn (Sequential)
STSTN/A (Reshaping)NoS-Attn, T-Attn (Parallel)
FG-STFormerS-TypeYesS-Attn, T-Attn (Sequential)
HyperformerS-TypeYesS-Attn, T-Conv (Sequential)
Human Interaction Recognition
IGFormerS-Type, T-TypeYesJoint-group-level ST-Attn
SkeleTRN/A (Pooling)YesJoint-group-level ST-Attn
ISTA-NetS-Type, T-TypeYesJoint-group-level ST-Attn
Both
SkateFormer4 Skate-TypesNoJoint-element-level ST-Attn

Partition, attend, reverse

SkateFormer framework. The input X of shape [T, V, C_in] passes three linear projection layers, and the Skate-Embedding of the frame indices t_idx is added. A SkateFormer block holds a self-attention part (layer norm, linear layer, channel-wise split into G-Conv, Skate-MSA and T-Conv branches, concatenation, linear layer, residual) and a feed-forward part (layer norm, FFN, residual). Four stages of two blocks each are separated by stride-2 temporal convolutions, from [T, V, C] to [T/8, V, 2C], followed by a head with pooling and a linear classifier that outputs y-hat.
The overall framework of SkateFormer. Linear projections and the Skate-Embedding feed eight SkateFormer blocks in four stages of two, with stride-2 temporal downsampling between stages; each block splits its channels into a graph convolution, a temporal convolution and the Skate-MSA, followed by a feed-forward layer.
  1. 01Four Skate-Types. Joints form neighboring partitions (limbs, torso) and distant partitions (same-position joints across them); frames form local segments of consecutive frames and N-strided global axes. Their four combinations are Skate-Type-1 to -4.
  2. 02Skate-MSA. Half of the channels go to self-attention, split into four parts that each attend within one Skate-Type partition. In the first blocks this is about 48× less computation than naive self-attention (38.87× measured).
  3. 03Skate-Embedding. The outer product of learnable skeletal features and fixed sinusoidal features of the sampled frame indices, trained with trimmed-uniform frame sampling and Bone Length AdaIN, which exchanges bone lengths between subjects.

BibTeX

@inproceedings{do2024skateformer,
  title={Skateformer: skeletal-temporal transformer for human action recognition},
  author={Do, Jeonghyeok and Kim, Munchurl},
  booktitle={European Conference on Computer Vision},
  pages={401--420},
  year={2024},
  organization={Springer}
}