Real-time character animation

MotionPersona

Real-Time Locomotion Control across Personas, Bodies, and Styles

Forty-four characters generated by one model, each conditioned on the persona and body shape of one captured participant, following the same straight path; the formation spreads because longer legs travel faster.

A generative locomotion controller with three controls: a captured motion persona (who is moving), a target body (an SMPL-X body shape that carries the motion) and a performed style (how the character is moving), under a trajectory command. One model covers 44 captured personas, a wide family of SMPL-X bodies and nine styles.

The controller is small: 36M parameters (175 MB), and it runs in real time on a CPU, at 27 ms per block on two threads of a laptop CPU. It is trained on thousands of hours of motion: MotionPersonaX holds 4,200 hours (33 hours of captured takes retargeted onto 128 bodies), and after the evaluation split is held out the VAE trains on 3,700 hours and the prior on 2,500 hours. Training is efficient: both stages finish in about 30 hours on consumer GPUs, under 100 GPU-hours in total (VAE: ~10 h on 2× RTX 4090; prior: ~19 h on 4× RTX 4090).

Overview video

01 · Controls

Change one dial, the other two stay put

Same command every time. Persona is a real performer's movement signature, body is the SMPL-X shape that carries it, style is how they move (happy, drunk, swimming…).

Top row: same person, five bodies, same rhythm. Bottom row: five people, same body, five different gaits.

02 · Demo

Run it yourself

Steer with the keyboard or a gamepad, switch persona and style, and drag ten body-shape sliders mid-walk. Raw model output, no foot IK. It runs locally: a Python install plus a free-for-research SMPL-X licence.

Steering with the keyboard
Switching the style (swimming)
Reshaping the body with the sliders while walking
Same persona and style on the new body

Get the demo

WASD move · Shift run · F lock facing · [ ] style · , . persona · B persona's own body. Run it with python realtime_demo.py (instructions). Rendering uses ai4animationpy; the SMPL-X model file is downloaded separately under its own license.

03 · Data

Capture, measure, then retarget

A performer is captured in exactly one body, so the raw data cannot say which traits belong to the person and which to the body. We build the missing cross-body data in three steps.

  1. 1Capture48 performers repeat the same nine styles and seven commands.
  2. 2MeasureFind which gait traits actually follow the body. Only trunk lean does.
  3. 3RetargetPut every take on 128 bodies, changing only what the measurement says.

1Capture

MotionPersona: about 33 hours of motion from 48 performers aged 5 to 68, recorded with optical motion capture. Everyone performs the same grid of nine styles × seven locomotion commands, so the effect of the person, of the style, and their interaction can be told apart. 44 performers also describe themselves with three persona attributes: role, affiliation and dominance.

2Measure

“The same motion on another body” has no unique answer, so we measure instead of guessing. Each of 27 gait descriptors is decomposed into a performer, a style and an interaction effect, and a trait is assigned to the body only if it survives an age control, holds in all nine styles, and is confirmed by twelve matched adult pairs that differ only in body.

Follows the body

Forward and side trunk lean grow with hip width. The retargeter uses 150° and 65° per metre, between the matched-pair and the population slopes.

Stays with the performer

Cadence, normalised step length, knee range, foot clearance, joint speed. A body-shape predictor explains at most 7.5 % of the differences between performers.

3Retarget

MotionPersonaX: every annotated clip optimised onto 128 SMPL-X bodies (the 48 captured ones plus 80 tiling heights of 1.05–1.95 m against eight girths), giving 328,960 clips, about 4,200 h. The optimiser keeps the source timing, makes trunk lean follow the body, and fixes penetration, foot contact and balance on the new body.

Follows the measurement. In the data, lean follows hip width (149.7°/m against the 150° target) while cadence and normalised step are identical on all three bodies.
Beats a plain copy. Watch the hand: the geometric copy (left) pushes it 9 cm into the torso; the optimisation (right) cuts that to 2.6 cm.

Body model, the eleven objective terms and the full measurement are in the MotionPersonaX dataset card.

04 · Explore

Browse the data in 3D

Any captured clip on the participant's SMPL-X body, and each participant's neutral walk retargeted onto any of the 128 bodies, streamed from the dataset and rendered in your browser.

Also available as a full-page Space; the data itself is on Hugging Face.

05 · Method

How it works

Two components, one job each: a decoder that knows the body, and a prior that knows the person.

Model overview: (a) shape-aware VAE, (b) latent flow-matching prior

Motion is generated block by block: every block predicts the next 45 frames (1.5 s at 30 fps) from the last frames already played, and only the first few of them are committed (12 in the offline protocol, 3 to 12 in the realtime demo) before the next block is generated, so the character reacts to new commands within a fraction of a second. Each block passes through two stages, trained one after the other.

Stage 1: shape-aware VAE

In panel (a), the encoder E compresses a 45-frame block x into a few latent tokens z (9 tokens of 48 dimensions). The decoder D turns z back into joint rotations, root motion and foot contacts on a given body: besides z it receives the body's 10 SMPL-X shape coefficients β and the last frames of the motion so far (xtail, the last 5 frames, so consecutive blocks join without a seam). It is trained with reconstruction, forward-kinematics and contact losses evaluated on the skeleton S(β) of that body.

The decoder therefore owns everything that is geometry: where the feet land, how far a step reaches, how the pelvis and trunk sit on this body. The latent z keeps what the motion is, and β decides how it is realised. Decoding one z sequence on different bodies gives the same motion carried by each of them:

One sequence of latent tokens (p21, happy, sampled by the prior on its own body) decoded on four SMPL-X bodies. Timing and gestures are shared; stride, speed (1.00 to 1.23 m/s) and posture follow the body.

Stage 2: latent flow-matching prior

In panel (b), with the codec frozen, a transformer (8 DiT blocks) learns to generate the next block's latent tokens z. It is conditioned on the motion history (one token from the codec encoder), the commanded trajectory (45 future positions and facings), the persona (a learned performer-ID token plus role, affiliation and dominance attribute tokens), the style, and the target body shape β. Sampling takes two Euler steps from Gaussian noise, after which the frozen decoder renders the tokens on the target body.

The prior chooses which motion this persona, in this style, would perform for this command on this body, in a space small enough for a real-time budget; the decoder fits the result to the body.

Streaming rollout: a new block every 12 committed frames. Timeline recorded on an RTX 5080 (9.9 ms per block against a 100 ms budget); on two threads of a laptop CPU a block takes 27 ms.

06 · Results

Experimental Results

The same gait descriptors from capture to generation, each control swept in isolation, compared against CAMDM, a state-of-the-art real-time diffusion controller trained on the same data.

1.0 %of frames with a foot sole more than 1 cm below the floor
(CAMDM: 49 %)
63 %of performers correctly identified from the generated motion (chance: 1 in 44)
(CAMDM: 32 %)
4.17 / 5match to the persona, 28 professionals
(CAMDM: 2.59)

Own-body test set, one-minute rollouts. CAMDM is marginally better on pose FPD (0.30 against our 0.32), and the two are level on foot smoothness and dynamics spread; ours wins on floor contact, persona and style.

Keeping the floor. The baseline loses the floor at both ends of the height range; ours keeps the feet where the data has them.
Not averaging people. The baseline averages two performers' pelvis bob; ours keeps part of the difference present in the data.
Beyond the performer's own body. One persona on its own body, on a training body it was seen on, on one it was withheld from, and on a body never seen in training.

07 · Extensions

Beyond locomotion

Both stages take their conditions as tokens, so a new task adds a token and retrains the prior. Two preliminary directions:

Character-aware in-betweening. Same keyframes on a 105 cm body: a method with no body input sinks the feet into the floor for the whole gap, ours keeps them where the data has them.
Speech-driven gesture. One speech clip of a BEAT2 speaker drives five other speakers on their own bodies, each in their own manner.

08 · Cite

BibTeX

@misc{shi2025motionpersonacharacteristicsawarelocomotioncontrol,
      title={MotionPersona: Characteristics-aware Locomotion Control},
      author={Mingyi Shi and Wei Liu and Jidong Mei and Wangpok Tse and Rui Chen and Xuelin Chen and Taku Komura},
      year={2025},
      eprint={2506.00173},
      archivePrefix={arXiv},
      primaryClass={cs.GR},
      url={https://arxiv.org/abs/2506.00173},
}