More research

MotionMaestro: Masked Tokenization for Unified Motion Generation

In the logo, “Mae” is marked as Masked AutoEncoder.

Yun Chen Munchurl Kim† Jeonghyeok Do†

Korea Advanced Institute of Science and Technology (KAIST), South Korea

† Co-corresponding authorscyruby@kaist.ac.krmkimee@kaist.ac.krehwjdgur0913@kaist.ac.kr

arXiv preprint, 2026

Abstract

Human motion generation plays an important role in applications such as character animation, virtual environments, and embodied interaction. While existing approaches have achieved remarkable progress, many of them are developed for individual tasks, including text-to-motion, pose-conditioned generation, and trajectory control. Although these tasks involve different types of conditions, a unified framework capable of handling them within a common representation would greatly simplify motion generation systems. We observe that diverse motion conditions can be naturally formulated as different observation patterns over motion sequences, where each task corresponds to a specific masking strategy.

Based on this insight, we introduce MotionMaestro, a unified motion generation framework that learns a shared representation for complete motions and heterogeneous partial observations through masked motion tokenization. MotionMaestro employs a three-stage training strategy that first learns a masked motion tokenizer, then refines its reconstruction ability on clean motions, and finally trains a conditional flow-matching generator in the learned latent space. Furthermore, we introduce an observation map and an observation loss to explicitly preserve provided motion conditions during generation. With this unified representation and conditioning mechanism, MotionMaestro supports text-guided and unconditional synthesis, pose conditioning and partial completion, temporal interpolation, trajectory control, and motion continuation. Experiments on the large-scale RoMo and MotionMillion datasets show state-of-the-art performance across diverse motion generation tasks.

One model, nine tasks

Each clip runs one motion through nine conditioning tasks with a single MotionMaestro model.

One MotionMaestro model per dataset; partial editing is zero-shot. Demo settings differ from the benchmark: keyframes every 32nd frame; the lower body is given for completion and editing.

Comparison with baselines

MDM, OmniControl, MotionLab and MotionMaestro on RoMo, next to the reference motion. All baselines are retrained with the same data splits.

the caption only

the man performs a push-up then stands up quickly.

MDM
OmniControlOmniCtrl
MotionLab
MotionMaestro (Ours)Ours
ReferenceRef.

The metric line under each column is part of the video. White labels mark baseline outputs that leave the camera view.

Frame interpolation with varying stride

Keyframes are given every 2nd to 32nd frame; one MotionMaestro model per dataset fills in the rest.

Paper results Tables 3 and 15

The tables report s = 2, 8, 16 and 32 (the main benchmark uses s = 4). The clips above are separate qualitative examples; a video panel does not correspond to a table cell.

Table 3. Frame interpolation with varying observation stride on RoMo.
Bold bestUnderline second best↓ lower is better
Scroll for more columns
Methods = 2s = 8s = 16s = 32
MPJPE  ↓FID  ↓MPJPE  ↓FID  ↓MPJPE  ↓FID  ↓MPJPE  ↓FID  ↓
MDM59.361.9252.2282.5375.2403.6431.3464.0
OmniControl803.2747.8929.2782.7943.3776.1930.6745.8
MotionLab11.110.718.914.135.421.457.530.4
MotionMaestro4.83.912.74.730.112.954.118.6
Table 15. Frame interpolation with varying observation stride on RoMo and MotionMillion. MPJPE (mm) is computed over unobserved frames; FID uses complete outputs. Lower is better; bold and underlined mark the best and second-best results per column.
Bold bestUnderline second best↓ lower is better
Scroll for more columns
MethodRoMoMotionMillion
s = 2s = 8s = 16s = 32s = 2s = 8s = 16s = 32
MPJPE  ↓FID  ↓MPJPE  ↓FID  ↓MPJPE  ↓FID  ↓MPJPE  ↓FID  ↓MPJPE  ↓FID  ↓MPJPE  ↓FID  ↓MPJPE  ↓FID  ↓MPJPE  ↓FID  ↓
MDM59.361.9252.2282.5375.2403.6431.3464.075.4128.7203.5301.0285.9401.1330.4454.8
OmniControl803.2747.8929.2782.7943.3776.1930.6745.8383.9503.6372.0473.5373.9457.7373.6426.0
MotionLab11.110.718.914.135.421.457.530.411.69.415.711.426.116.947.027.8
MotionMaestro4.83.912.74.730.112.954.118.62.81.57.51.820.35.943.115.2
MotionMaestro-5B4.34.311.14.426.010.147.815.62.31.06.31.615.94.536.011.4

Partial completion from different body parts

Only the torso, the arms or the legs are given at every frame; the model completes the whole body.

Paper results Tables 4 and 16

The tables use torso, head+hands and one arm on RoMo, and upper, head+hands and one arm on MotionMillion. The clips use Torso / Arms / Legs presets; a video panel does not correspond to a table column.

Table 4. Partial completion under different observed body parts on RoMo.
Bold bestUnderline second best↓ lower is better
Scroll for more columns
Methodtorsohead+handsone arm
MPJPE  ↓FID  ↓MPJPE  ↓FID  ↓MPJPE  ↓FID  ↓
MDM301.4553.8346.0516.8313.5561.4
OmniControl535.3477.5627.3631.5297.4293.1
MotionLab378.1379.373.835.381.232.7
MotionMaestro109.344.768.531.879.539.0
Table 16. Partial completion under different observed body parts on RoMo and MotionMillion. We report MPJPE (mm) over unobserved joints and FID. Observed body parts are specified separately for each dataset. Lower is better; bold and underlined mark the best and second-best results, respectively.
Bold bestUnderline second best↓ lower is better
Scroll for more columns
MethodRoMoMotionMillion
torsohead+handsone armupperhead+handsone arm
MPJPE  ↓FID  ↓MPJPE  ↓FID  ↓MPJPE  ↓FID  ↓MPJPE  ↓FID  ↓MPJPE  ↓FID  ↓MPJPE  ↓FID  ↓
MDM301.4553.8346.0516.8313.5561.4206.8401.9247.9469.3230.6470.0
OmniControl535.3477.5627.3631.5297.4293.1349.6448.5349.6448.5252.5287.0
MotionLab378.1379.373.835.381.232.771.124.368.647.774.740.5
MotionMaestro109.344.768.531.879.539.061.323.573.767.079.762.2
MotionMaestro-5B95.842.256.730.368.233.645.213.860.646.464.646.3

Paper figures

MotionMaestro better preserves the given keyframes than MDM and MotionLab, and retains the given body parts while completing the missing motion.

Partial editing is zero-shot and preliminary; instruction adherence varies across examples.

Quantitative results

Best on all ten tasks on RoMo and on eight of ten on MotionMillion, with one generator checkpoint per dataset.

Table 1. Unified motion generation on RoMo.
Bold bestUnderline second best↓ lower is better
Scroll for more columns
MethodVenueUncond.T2MP2MTP2MFLFFIExtrap.PartialTrajectoryLong
FID  ↓FID  ↓FID  ↓FID  ↓FID  ↓MPJPE  ↓FID  ↓MPJPE  ↓FID  ↓FID  ↓
MDMICLR2023507.6450.7518.4456.1523.3117.6342.8278.9–510.9
OmniControlICLR2024429.1199.4771.0312.8753.2875.9740.3627.3283.7398.6
MotionLabICCV202576.7101.553.546.530.012.728.171.889.3123.5
MotionMaestro (Ours)–20.414.934.828.720.86.719.461.059.226.2

Each method is evaluated on its supported tasks. FI and Partial report MPJPE (mm) over unobserved frames and joints, respectively; the other tasks report FID, computed with the frozen MotionMillion evaluator, which is out of domain on RoMo. Baselines are retrained on each dataset with the same splits; compute budgets are not matched.

Ablations Tables 5, 9 and 11
Table 5. Ablations of tokenizer training and observation conditioning on RoMo. The upper block compares training on clean sequences only, random masking, and structured masking (MotionMaestro, bottom row). The lower block uses the same tokenizer trained with structured masking and separately removes the generator’s observation map input Oc or observation loss ℒobs.
Bold bestUnderline second best↓ lower is better
Scroll for more columns
MethodUncond.T2MP2MTP2MFLFFIExtrap.PartialTrajectoryLong
FID  ↓FID  ↓FID  ↓FID  ↓FID  ↓MPJPE  ↓FID  ↓MPJPE  ↓FID  ↓FID  ↓
Tokenizer training (Stage 1)
Clean sequences only35.518.338.331.024.27.822.161.872.941.2
Random masking22.315.637.830.621.07.521.363.759.736.2
Generator training (Stage 3)
w/o observation map Oc18.615.036.929.620.96.219.661.357.530.9
w/o ℒobs21.815.235.929.820.16.819.961.158.928.7
MotionMaestro (Ours)20.414.934.828.720.86.719.461.059.226.2
Table 9. Masked-motion reconstruction after Stage 2. All tokenizers are evaluated on the same four masking families using their refined decoders, without the latent generator. We report MPJPE (mm) over hidden regions only. Lower is better; bold and underlined values denote the best and second-best results, respectively.
Bold bestUnderline second best↓ lower is better
Scroll for more columns
TokenizerAnatomical tube ↓Temporal stride ↓Feature-group ↓Mixed ↓
Clean-only AE257.8207.4273.1249.4
Random-mask AE201.79.3306.7199.8
Structured-mask AE (Ours)46.27.0158.743.8
Table 11. Observation preservation on RoMo. The first six columns report observed-joint MPJPE; Trajectory reports observed root-xz error. All values are in millimeters on 512 clips per task. All variants use the same frozen tokenizer. Bold and underlined values mark the lowest and second-lowest point estimates, not statistical significance.
Bold lowestUnderline second lowest↓ lower is better
Scroll for more columns
VariantFLF ↓FI ↓Extrap. ↓Partial ↓P2M ↓TP2M ↓Trajectory ↓
w/o observation map5.504.322.923.566.255.928.06
w/o ℒobs5.783.782.983.746.405.947.93
MotionMaestro (Ours)4.853.702.913.495.064.637.51
MotionMaestro-5B and computational cost Tables 13 and 14

MotionMaestro-5B also uses pooled CLIP text conditioning, so these results characterize the larger configuration rather than isolating parameter count.

Table 13. Unified motion generation including MotionMaestro-5B on RoMo. FI and Partial report MPJPE (mm) over unobserved frames and joints, respectively; the remaining tasks report FID. Lower is better; bold and underlined mark the best and second-best results per task.
Bold bestUnderline second best↓ lower is better
Scroll for more columns
MethodVenueUncond.T2MP2MTP2MFLFFIExtrap.PartialTrajectoryLong
FID  ↓FID  ↓FID  ↓FID  ↓FID  ↓MPJPE  ↓FID  ↓MPJPE  ↓FID  ↓FID  ↓
MDMICLR2023507.6450.7518.4456.1523.3117.6342.8278.9–510.9
OmniControlICLR2024429.1199.4771.0312.8753.2875.9740.3627.3283.7398.6
MotionLabICCV202576.7101.553.546.530.012.728.171.889.3123.5
MotionMaestro (Ours)–20.414.934.828.720.86.719.461.059.226.2
MotionMaestro-5B (Ours)–20.317.231.328.618.45.817.947.259.019.0
Table 14. Unified motion generation and computational cost on MotionMillion. Evaluation settings, metrics, and notation follow Table 13. We additionally report computational cost.
Bold bestUnderline second best↓ lower is better
Scroll for more columns
MethodUncond.T2MP2MTP2MFLFFIExtrap.PartialTrajectoryLongInferenceMemoryParams
FID  ↓FID  ↓FID  ↓FID  ↓FID  ↓MPJPE  ↓FID  ↓MPJPE  ↓FID  ↓FID  ↓Time (s)GBBillion (B)
MDM427.7297.6523.8334.4530.6122.8370.4206.8–336.24.690.40.081
OmniControl376.3153.9366.8193.8365.5375.4551.3349.6276.6177.961.580.50.102
MotionLab135.161.474.449.450.712.338.871.194.4108.60.972.60.367
MotionMaestro67.236.274.352.844.53.836.961.3118.246.41.2811.71.285
MotionMaestro-5B67.934.751.342.832.53.224.345.2109.535.81.3130.65.096
What each task gives the model Table 21
Table 21. The ten main-paper evaluation tasks. A pose includes joint positions and rotations together with root placement and heading. Text is provided only for T2M, TP2M, and Long. MPJPE is evaluated on unobserved joint positions; FID uses the whole output sequence.
Scroll for more columns
TaskMotion observationTextMain metric
Uncond.NoneNoFID
T2MNoneYesFID
P2MFirst poseNoFID
TP2MFirst poseYesFID
FLFFirst and last posesNoFID
FIPeriodic poses at spacing s = 4NoHidden MPJPE
Extrap.A motion prefixNoFID
PartialPositions of 13 upper-body joints at all framesNoHidden MPJPE
TrajectoryRoot xz at all frames; no heading or joint posesNoFID
LongNone; 129–300 output framesYesFID

The MAE in Maestro

Training pipeline. Stage 1: masked motion and its observation map pass through the motion embedding encoder, encoder, decoder and motion embedding decoder, trained with the tokenizer loss against the clean motion. Stage 2: the encoder side is frozen and only the decoder side is fine-tuned on clean motion. Stage 3: the frozen encoder embeds the target and the observed condition; the noisy target latent, the condition latent and the observation map are concatenated and fed to a transformer generator with double-stream and single-stream blocks and a text condition, trained with the flow matching loss and an observation loss on the decoded prediction.
Overview of MotionMaestro. Stages 1–2 train the masked tokenizer and refine its decoder; Stage 3 trains the flow-matching generator.
  1. 01Masked motion tokenizer. Inspired by MAETok, the tokenizer is first trained with task-aligned masking, then only its decoder is refined on clean motions.
  2. 02Shared latent space. One latent space for complete motions and partial observations.
  3. 03Conditional flow matching. An observation map marks what is given; an observation loss encourages outputs to match it.

BibTeX

@article{chen2026motionmaestro,
  title={MotionMaestro: Masked Tokenization for Unified Motion Generation},
  author={Chen, Yun and Kim, Munchurl and Do, Jeonghyeok},
  journal={arXiv preprint arXiv:XXXX.XXXXX},
  year={2026}
}