One model, nine tasks
Each clip runs one motion through nine conditioning tasks with a single MotionMaestro model.
One MotionMaestro model per dataset; partial editing is zero-shot. Demo settings differ from the benchmark: keyframes every 32nd frame; the lower body is given for completion and editing.
Comparison with baselines
MDM, OmniControl, MotionLab and MotionMaestro on RoMo, next to the reference motion. All baselines are retrained with the same data splits.
the caption only
the man performs a push-up then stands up quickly.
the woman dances forward with rhythmic arm swings and steps.
the first poseThe metric line burned into these clips is labelled “I2M FID”; it is the paper’s P2M FID.
the first pose and the captionThe metric line burned into these clips is labelled “TI2M FID”; it is the paper’s TP2M FID.
the woman stands on stage and gestures with her hands while talking to the audience.
the man stands on a slackline, moving his arms up and down to maintain his balance.
the first and last pose
every 4th frame
the first half of the clipThe numbers burned into these two clips are an extrapolation MPJPE, which the paper does not report. The paper evaluates extrapolation with FID (Tables 1 and 2).
the upper body of every frame (no whole pose is ever given)
the root path only (orange), no pose at any frameMDM cannot directly receive an absolute root-xz trajectory, so its panel is marked “not applicable”.
Same source clip as First-Last Frame example 1.
The metric line under each column is part of the video. White labels mark baseline outputs that leave the camera view.
Frame interpolation with varying stride
Keyframes are given every 2nd to 32nd frame; one MotionMaestro model per dataset fills in the rest.
Paper results Tables 3 and 15
The tables report s = 2, 8, 16 and 32 (the main benchmark uses s = 4). The clips above are separate qualitative examples; a video panel does not correspond to a table cell.
| Method | s = 2 | s = 8 | s = 16 | s = 32 | ||||
|---|---|---|---|---|---|---|---|---|
| MPJPE ↓ | FID ↓ | MPJPE ↓ | FID ↓ | MPJPE ↓ | FID ↓ | MPJPE ↓ | FID ↓ | |
| MDM | 59.3 | 61.9 | 252.2 | 282.5 | 375.2 | 403.6 | 431.3 | 464.0 |
| OmniControl | 803.2 | 747.8 | 929.2 | 782.7 | 943.3 | 776.1 | 930.6 | 745.8 |
| MotionLab | 11.1 | 10.7 | 18.9 | 14.1 | 35.4 | 21.4 | 57.5 | 30.4 |
| MotionMaestro | 4.8 | 3.9 | 12.7 | 4.7 | 30.1 | 12.9 | 54.1 | 18.6 |
| Method | RoMo | MotionMillion | ||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| s = 2 | s = 8 | s = 16 | s = 32 | s = 2 | s = 8 | s = 16 | s = 32 | |||||||||
| MPJPE ↓ | FID ↓ | MPJPE ↓ | FID ↓ | MPJPE ↓ | FID ↓ | MPJPE ↓ | FID ↓ | MPJPE ↓ | FID ↓ | MPJPE ↓ | FID ↓ | MPJPE ↓ | FID ↓ | MPJPE ↓ | FID ↓ | |
| MDM | 59.3 | 61.9 | 252.2 | 282.5 | 375.2 | 403.6 | 431.3 | 464.0 | 75.4 | 128.7 | 203.5 | 301.0 | 285.9 | 401.1 | 330.4 | 454.8 |
| OmniControl | 803.2 | 747.8 | 929.2 | 782.7 | 943.3 | 776.1 | 930.6 | 745.8 | 383.9 | 503.6 | 372.0 | 473.5 | 373.9 | 457.7 | 373.6 | 426.0 |
| MotionLab | 11.1 | 10.7 | 18.9 | 14.1 | 35.4 | 21.4 | 57.5 | 30.4 | 11.6 | 9.4 | 15.7 | 11.4 | 26.1 | 16.9 | 47.0 | 27.8 |
| MotionMaestro | 4.8 | 3.9 | 12.7 | 4.7 | 30.1 | 12.9 | 54.1 | 18.6 | 2.8 | 1.5 | 7.5 | 1.8 | 20.3 | 5.9 | 43.1 | 15.2 |
| MotionMaestro-5B | 4.3 | 4.3 | 11.1 | 4.4 | 26.0 | 10.1 | 47.8 | 15.6 | 2.3 | 1.0 | 6.3 | 1.6 | 15.9 | 4.5 | 36.0 | 11.4 |
Partial completion from different body parts
Only the torso, the arms or the legs are given at every frame; the model completes the whole body.
Paper results Tables 4 and 16
The tables use torso, head+hands and one arm on RoMo, and upper, head+hands and one arm on MotionMillion. The clips use Torso / Arms / Legs presets; a video panel does not correspond to a table column.
| Method | torso | head+hands | one arm | |||
|---|---|---|---|---|---|---|
| MPJPE ↓ | FID ↓ | MPJPE ↓ | FID ↓ | MPJPE ↓ | FID ↓ | |
| MDM | 301.4 | 553.8 | 346.0 | 516.8 | 313.5 | 561.4 |
| OmniControl | 535.3 | 477.5 | 627.3 | 631.5 | 297.4 | 293.1 |
| MotionLab | 378.1 | 379.3 | 73.8 | 35.3 | 81.2 | 32.7 |
| MotionMaestro | 109.3 | 44.7 | 68.5 | 31.8 | 79.5 | 39.0 |
| Method | RoMo | MotionMillion | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| torso | head+hands | one arm | upper | head+hands | one arm | |||||||
| MPJPE ↓ | FID ↓ | MPJPE ↓ | FID ↓ | MPJPE ↓ | FID ↓ | MPJPE ↓ | FID ↓ | MPJPE ↓ | FID ↓ | MPJPE ↓ | FID ↓ | |
| MDM | 301.4 | 553.8 | 346.0 | 516.8 | 313.5 | 561.4 | 206.8 | 401.9 | 247.9 | 469.3 | 230.6 | 470.0 |
| OmniControl | 535.3 | 477.5 | 627.3 | 631.5 | 297.4 | 293.1 | 349.6 | 448.5 | 349.6 | 448.5 | 252.5 | 287.0 |
| MotionLab | 378.1 | 379.3 | 73.8 | 35.3 | 81.2 | 32.7 | 71.1 | 24.3 | 68.6 | 47.7 | 74.7 | 40.5 |
| MotionMaestro | 109.3 | 44.7 | 68.5 | 31.8 | 79.5 | 39.0 | 61.3 | 23.5 | 73.7 | 67.0 | 79.7 | 62.2 |
| MotionMaestro-5B | 95.8 | 42.2 | 56.7 | 30.3 | 68.2 | 33.6 | 45.2 | 13.8 | 60.6 | 46.4 | 64.6 | 46.3 |
Paper figures
MotionMaestro better preserves the given keyframes than MDM and MotionLab, and retains the given body parts while completing the missing motion.
Partial editing is zero-shot and preliminary; instruction adherence varies across examples.
Quantitative results
Best on all ten tasks on RoMo and on eight of ten on MotionMillion, with one generator checkpoint per dataset.
| Method | Venue | Uncond. | T2M | P2M | TP2M | FLF | FI | Extrap. | Partial | Trajectory | Long |
|---|---|---|---|---|---|---|---|---|---|---|---|
| FID ↓ | FID ↓ | FID ↓ | FID ↓ | FID ↓ | MPJPE ↓ | FID ↓ | MPJPE ↓ | FID ↓ | FID ↓ | ||
| MDM | ICLR2023 | 507.6 | 450.7 | 518.4 | 456.1 | 523.3 | 117.6 | 342.8 | 278.9 | – | 510.9 |
| OmniControl | ICLR2024 | 429.1 | 199.4 | 771.0 | 312.8 | 753.2 | 875.9 | 740.3 | 627.3 | 283.7 | 398.6 |
| MotionLab | ICCV2025 | 76.7 | 101.5 | 53.5 | 46.5 | 30.0 | 12.7 | 28.1 | 71.8 | 89.3 | 123.5 |
| MotionMaestro (Ours) | – | 20.4 | 14.9 | 34.8 | 28.7 | 20.8 | 6.7 | 19.4 | 61.0 | 59.2 | 26.2 |
| Method | Uncond. | T2M | P2M | TP2M | FLF | FI | Extrap. | Partial | Trajectory | Long | Inference | Memory | Params |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| FID ↓ | FID ↓ | FID ↓ | FID ↓ | FID ↓ | MPJPE ↓ | FID ↓ | MPJPE ↓ | FID ↓ | FID ↓ | Time (s) | GB | Billion (B) | |
| MDM | 427.7 | 297.6 | 523.8 | 334.4 | 530.6 | 122.8 | 370.4 | 206.8 | – | 336.2 | 4.69 | 0.4 | 0.081 |
| OmniControl | 376.3 | 153.9 | 366.8 | 193.8 | 365.5 | 375.4 | 551.3 | 349.6 | 276.6 | 177.9 | 61.58 | 0.5 | 0.102 |
| MotionLab | 135.1 | 61.4 | 74.4 | 49.4 | 50.7 | 12.3 | 38.8 | 71.1 | 94.4 | 108.6 | 0.97 | 2.6 | 0.367 |
| MotionMaestro | 67.2 | 36.2 | 74.3 | 52.8 | 44.5 | 3.8 | 36.9 | 61.3 | 118.2 | 46.4 | 1.28 | 11.7 | 1.285 |
Each method is evaluated on its supported tasks. FI and Partial report MPJPE (mm) over unobserved frames and joints, respectively; the other tasks report FID, computed with the frozen MotionMillion evaluator, which is out of domain on RoMo. Baselines are retrained on each dataset with the same splits; compute budgets are not matched.
Ablations Tables 5, 9 and 11
| Method | Uncond. | T2M | P2M | TP2M | FLF | FI | Extrap. | Partial | Trajectory | Long |
|---|---|---|---|---|---|---|---|---|---|---|
| FID ↓ | FID ↓ | FID ↓ | FID ↓ | FID ↓ | MPJPE ↓ | FID ↓ | MPJPE ↓ | FID ↓ | FID ↓ | |
| Tokenizer training (Stage 1) | ||||||||||
| Clean sequences only | 35.5 | 18.3 | 38.3 | 31.0 | 24.2 | 7.8 | 22.1 | 61.8 | 72.9 | 41.2 |
| Random masking | 22.3 | 15.6 | 37.8 | 30.6 | 21.0 | 7.5 | 21.3 | 63.7 | 59.7 | 36.2 |
| Generator training (Stage 3) | ||||||||||
| w/o observation map Oc | 18.6 | 15.0 | 36.9 | 29.6 | 20.9 | 6.2 | 19.6 | 61.3 | 57.5 | 30.9 |
| w/o ℒobs | 21.8 | 15.2 | 35.9 | 29.8 | 20.1 | 6.8 | 19.9 | 61.1 | 58.9 | 28.7 |
| MotionMaestro (Ours) | 20.4 | 14.9 | 34.8 | 28.7 | 20.8 | 6.7 | 19.4 | 61.0 | 59.2 | 26.2 |
| Tokenizer | Anatomical tube ↓ | Temporal stride ↓ | Feature-group ↓ | Mixed ↓ |
|---|---|---|---|---|
| Clean-only AE | 257.8 | 207.4 | 273.1 | 249.4 |
| Random-mask AE | 201.7 | 9.3 | 306.7 | 199.8 |
| Structured-mask AE (Ours) | 46.2 | 7.0 | 158.7 | 43.8 |
| Variant | FLF ↓ | FI ↓ | Extrap. ↓ | Partial ↓ | P2M ↓ | TP2M ↓ | Trajectory ↓ |
|---|---|---|---|---|---|---|---|
| w/o observation map | 5.50 | 4.32 | 2.92 | 3.56 | 6.25 | 5.92 | 8.06 |
| w/o ℒobs | 5.78 | 3.78 | 2.98 | 3.74 | 6.40 | 5.94 | 7.93 |
| MotionMaestro (Ours) | 4.85 | 3.70 | 2.91 | 3.49 | 5.06 | 4.63 | 7.51 |
MotionMaestro-5B and computational cost Tables 13 and 14
MotionMaestro-5B also uses pooled CLIP text conditioning, so these results characterize the larger configuration rather than isolating parameter count.
| Method | Venue | Uncond. | T2M | P2M | TP2M | FLF | FI | Extrap. | Partial | Trajectory | Long |
|---|---|---|---|---|---|---|---|---|---|---|---|
| FID ↓ | FID ↓ | FID ↓ | FID ↓ | FID ↓ | MPJPE ↓ | FID ↓ | MPJPE ↓ | FID ↓ | FID ↓ | ||
| MDM | ICLR2023 | 507.6 | 450.7 | 518.4 | 456.1 | 523.3 | 117.6 | 342.8 | 278.9 | – | 510.9 |
| OmniControl | ICLR2024 | 429.1 | 199.4 | 771.0 | 312.8 | 753.2 | 875.9 | 740.3 | 627.3 | 283.7 | 398.6 |
| MotionLab | ICCV2025 | 76.7 | 101.5 | 53.5 | 46.5 | 30.0 | 12.7 | 28.1 | 71.8 | 89.3 | 123.5 |
| MotionMaestro (Ours) | – | 20.4 | 14.9 | 34.8 | 28.7 | 20.8 | 6.7 | 19.4 | 61.0 | 59.2 | 26.2 |
| MotionMaestro-5B (Ours) | – | 20.3 | 17.2 | 31.3 | 28.6 | 18.4 | 5.8 | 17.9 | 47.2 | 59.0 | 19.0 |
| Method | Uncond. | T2M | P2M | TP2M | FLF | FI | Extrap. | Partial | Trajectory | Long | Inference | Memory | Params |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| FID ↓ | FID ↓ | FID ↓ | FID ↓ | FID ↓ | MPJPE ↓ | FID ↓ | MPJPE ↓ | FID ↓ | FID ↓ | Time (s) | GB | Billion (B) | |
| MDM | 427.7 | 297.6 | 523.8 | 334.4 | 530.6 | 122.8 | 370.4 | 206.8 | – | 336.2 | 4.69 | 0.4 | 0.081 |
| OmniControl | 376.3 | 153.9 | 366.8 | 193.8 | 365.5 | 375.4 | 551.3 | 349.6 | 276.6 | 177.9 | 61.58 | 0.5 | 0.102 |
| MotionLab | 135.1 | 61.4 | 74.4 | 49.4 | 50.7 | 12.3 | 38.8 | 71.1 | 94.4 | 108.6 | 0.97 | 2.6 | 0.367 |
| MotionMaestro | 67.2 | 36.2 | 74.3 | 52.8 | 44.5 | 3.8 | 36.9 | 61.3 | 118.2 | 46.4 | 1.28 | 11.7 | 1.285 |
| MotionMaestro-5B | 67.9 | 34.7 | 51.3 | 42.8 | 32.5 | 3.2 | 24.3 | 45.2 | 109.5 | 35.8 | 1.31 | 30.6 | 5.096 |
What each task gives the model Table 21
| Task | Motion observation | Text | Main metric |
|---|---|---|---|
| Uncond. | None | No | FID |
| T2M | None | Yes | FID |
| P2M | First pose | No | FID |
| TP2M | First pose | Yes | FID |
| FLF | First and last poses | No | FID |
| FI | Periodic poses at spacing s = 4 | No | Hidden MPJPE |
| Extrap. | A motion prefix | No | FID |
| Partial | Positions of 13 upper-body joints at all frames | No | Hidden MPJPE |
| Trajectory | Root xz at all frames; no heading or joint poses | No | FID |
| Long | None; 129–300 output frames | Yes | FID |
The MAE in Maestro
- 01Masked motion tokenizer. Inspired by MAETok, the tokenizer is first trained with task-aligned masking, then only its decoder is refined on clean motions.
- 02Shared latent space. One latent space for complete motions and partial observations.
- 03Conditional flow matching. An observation map marks what is given; an observation loss encourages outputs to match it.
BibTeX
@article{chen2026motionmaestro,
title={MotionMaestro: Masked Tokenization for Unified Motion Generation},
author={Chen, Yun and Kim, Munchurl and Do, Jeonghyeok},
journal={arXiv preprint arXiv:XXXX.XXXXX},
year={2026}
}