Part-factorized conditioning
Content and style conditions retain their anatomical indices, with the root trajectory represented separately.
Department of Computer Science, College of AI, Cyber and Computing
The University of Texas at San Antonio
The Visual Computer Under Review
Editing character motion often requires borrowing an arm gesture or leg gait from reference sequences while retaining the original action, timing, root path, and unselected body regions. Motion datasets rarely provide paired targets for arbitrary local content–reference combinations, and self-reconstruction can let a diffusion model copy content without learning to respond to local references. We present MoSAIC for controllable character animation. The latent diffusion framework factorizes content and style by body part, conditions root trajectory separately, and uses a user-specified map to route references to anatomical regions. To make this behavior learnable, aligned intervention supervision applies a known local transformation to a source motion and constructs synchronized references and counterfactual targets that distinguish the requested response from what should remain unchanged. Anatomically constrained injection then supplies the routed cues during denoising. The validation-selected model retains HumanML3D generation quality. In a frozen evaluation of 128 motions and 896 routed conditions, masked routing achieves a fractional selected-target improvement of 0.198. Compared with whole-body routing, it reduces preserved-region error from 70.64 to 66.45 mm and matched-noise off-target leakage from 18.08 to 9.88 mm. MoSAIC therefore provides a more favorable response–preservation trade-off for part-local editing. Arm responses remain weaker than leg responses, and quantitative population-level natural-reference and multi-reference generalization remain unevaluated.
Content and style conditions retain their anatomical indices, with the root trajectory represented separately.
A source map selects content or a reference motion independently for each anatomical region.
Known local transformations create synchronized references and counterfactual targets, identifying both the requested response and the motion that should be preserved.
Reference cues are routed to selected regions while unselected regions receive content-derived conditions, limiting off-target influence.
On the frozen evaluation of 128 motions and 896 routed conditions, masked routing improves the balance between local response and preservation:
Arms remain weaker than legs and spine. These results show a better response–preservation balance; they do not establish overall state-of-the-art motion style transfer.
Please cite the arXiv preprint below.
@misc{amini2026mosaic,
title = {MoSAIC: Aligned Intervention Supervision
for Part-Local Motion Style Transfer},
author = {Amini, Nazanin and Desai, Kevin},
year = {2026},
eprint = {2607.26304},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
url = {https://arxiv.org/abs/2607.26304}
}
This work was partially supported by the U.S. National Science Foundation under Award Nos. 2211785 and 2316240.