Complementary to REPA
At 100K steps, combining REPI with REPA achieves an FID of 11.78, outperforming REPI (14.51) and REPA (19.40) alone.
1 Sun Yat-sen University2 Video Rebirth3 The Hong Kong Polytechnic University
At 100K steps, combining REPI with REPA achieves an FID of 11.78, outperforming REPI (14.51) and REPA (19.40) alone.
REPI + REPA reaches an FID of 8.22 in only 160K training steps, matching vanilla SiT trained for 7M steps (FID 8.30), corresponding to over 43.5× fewer training steps.
The encoder and projection layers are used only during training and are fully removed at inference, leaving the original diffusion backbone unchanged and introducing no additional inference cost.
Recent representation alignment (REPA) methods accelerate diffusion transformer training by aligning projections of the transformer's hidden states with representations from pretrained visual encoders. In this work, we explore a reverse and complementary direction to REPA: rather than projecting diffusion representations into the encoder's space, we inject encoder representations into the diffusion transformer, allowing them to actively participate in the denoising process. To this end, we introduce REPresentation Injection (REPI), a training framework based on a scaffold-to-internalization strategy, in which projected encoder representations initially serve as a temporary scaffold and are then progressively internalized by the diffusion transformer. REPI outperforms REPA across a wide range of backbones and is highly complementary to it: combining the two yields substantial gains over either alone. Notably, with only 160K training steps, REPI + REPA matches vanilla SiT trained for 7M steps, a speedup of over 43.5×.
Early training
Projected encoder keys and values temporarily replace the model's native K/V, while the model's native queries preserve conditioning on the noisy input.
Subsequent training
The model resumes computing its own K/V, while an internalization objective encourages the native representations to match the projected encoder targets.
The visual encoder and projection layers are removed. Sampling is performed using only the original diffusion backbone.
REPI not only outperforms REPA across a wide range of backbones but is also highly complementary to it, and combining the two yields substantial gains over either alone.
REPI generalizes across visual encoders, diffusion backbones, and image resolutions.
The temporary scaffold alone yields substantial improvements that persist after its removal. Adding the internalization objective further improves generation quality.
@misc{fu2026scaffold,
title = {Scaffold Then Internalize: Representation Injection for Diffusion Transformers},
author = {Han Fu and Jiacheng Chen and Baoquan Zhao and Weidong Chen and Wei Liu and Qing Li and Xudong Mao},
year = {2026},
eprint = {2609.35292},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
url = {https://arxiv.org/abs/2609.35292}
}