One-to-All Animation: Alignment-Free
Character Animation and Image Pose Transfer

MY ALT TEXT

We introduce One-to-All Animation, a unified framework for pose-driven personalized generation. Unlike prior methods that require both spatially-aligned references and pose retargeting, our framework supports: (1) cross-scale video animation with either retargeted or original driving motion, (2) cross-scale image pose transfer, and (3) temporally coherent long video generation.

Abstract

Recent advances in diffusion models have greatly improved pose-driven character animation. However, existing methods are limited to spatially aligned reference-pose pairs with matched skeletal structures. Handling reference-pose misalignment remains unsolved. To address this, we present One-to-All Animation, a unified framework for high-fidelity character animation and image pose transfer for references with arbitrary layouts. First, to handle spatially misaligned reference, we reformulate training as a self-supervised outpainting task that transforms diverse-layout reference into a unified occluded-input format. Second, to process partially visible reference, we design a reference extractor for comprehensive identity feature extraction. Further, we integrate hybrid reference fusion attention to handle varying resolutions and dynamic sequence lengths. Finally, from the perspective of generation quality, we introduce identity-robust pose control that decouples appearance from skeletal structure to mitigate pose overfitting, and a token replace strategy for coherent long-video generation. Extensive experiments show that our method outperforms existing approaches. The code and model will be available at https://github.com/ssj9596/One-to-All-Animation.

Method

MY ALT TEXT

Overview of the proposed framework. We introduce outpainting preprocess to handle diverse body proportions through face-centered random masking during training and pose-guided translation at inference. The driving poses are encoded and refined via reference-guided pose control to preserve facial identity despite skeletal mismatch. Reference features are progressively injected through hybrid reference fusion attention, supporting variable resolutions and dynamic sequence lengths.

place gif

Generated Videos



Comparison with SOTA

Aligned Animation

Misaligned Animation

Reference
Reference
Input
True Input
Output Video
Reference
Reference
Input
True Input
Output Video
Reference
Reference
Input
True Input
Output Video
Reference
Reference
Input
True Input
Output Video
Reference
Reference
Input
True Input
Output Video

Text Edits of Unseen Elements

a clown in a red suit dancing joyfully on a street.
a clown in a red suit and blue pants dancing joyfully on a street.
a clown in a red suit dancing joyfully on a street, with a McDonald's sign visible in the background.