Scaffold Then Internalize:
Representation Injection for Diffusion Transformers

1 Sun Yat-sen University2 Video Rebirth3 The Hong Kong Polytechnic University

Samples from REPI + REPA on SiT-XL/2 at 512 × 512, trained for 400K steps with CFG = 4.0.

Complementary to REPA

At 100K steps, combining REPI with REPA achieves an FID of 11.78, outperforming REPI (14.51) and REPA (19.40) alone.

160K vs. 7M

REPI + REPA reaches an FID of 8.22 in only 160K training steps, matching vanilla SiT trained for 7M steps (FID 8.30), corresponding to over 43.5× fewer training steps.

Original Backbone at Inference

The encoder and projection layers are used only during training and are fully removed at inference, leaving the original diffusion backbone unchanged and introducing no additional inference cost.

Abstract

Recent representation alignment (REPA) methods accelerate diffusion transformer training by aligning projections of the transformer's hidden states with representations from pretrained visual encoders. In this work, we explore a reverse and complementary direction to REPA: rather than projecting diffusion representations into the encoder's space, we inject encoder representations into the diffusion transformer, allowing them to actively participate in the denoising process. To this end, we introduce REPresentation Injection (REPI), a training framework based on a scaffold-to-internalization strategy, in which projected encoder representations initially serve as a temporary scaffold and are then progressively internalized by the diffusion transformer. REPI outperforms REPA across a wide range of backbones and is highly complementary to it: combining the two yields substantial gains over either alone. Notably, with only 160K training steps, REPI + REPA matches vanilla SiT trained for 7M steps, a speedup of over 43.5×.

Method Overview

01

Early training

Scaffold

Frozen encoder keys and values are projected into a diffusion transformer attention layer; native queries and the residual path remain.

Projected encoder keys and values temporarily replace the model's native K/V, while the model's native queries preserve conditioning on the noisy input.

02

Subsequent training

Internalization

Native keys and values are restored and aligned with the projected encoder representations using an internalization objective.

The model resumes computing its own K/V, while an internalization objective encourages the native representations to match the projected encoder targets.

03

Inference

The visual encoder and projection layers are removed. Sampling is performed using only the original diffusion backbone.

Inference

Quantitative Results

REPI not only outperforms REPA across a wide range of backbones but is also highly complementary to it, and combining the two yields substantial gains over either alone.

FID curves for SiT-B, SiT-L, and SiT-XL at 100K, 200K, and 400K steps; REPI improves upon REPA and their combination improves further.
ImageNet 256 × 256 · without CFG.
On SiT-XL/2, REPI + REPA trained for 200K steps already outperforms REPA trained for 400K steps.

Qualitative Results

iREPA and REPI are evaluated using the same initial noise and sampler, without classifier-free guidance.

Three groups of generated samples at 100K, 200K, and 400K training steps. iREPA is the top row and REPI is the bottom row.
SiT-XL/2. Top: iREPA. Bottom: REPI. Each column uses the same initial noise and shows results after 100K, 200K, and 400K training steps.

Generality

REPI generalizes across visual encoders, diffusion backbones, and image resolutions.

Visual Encoders

REPI outperforms REPA across all ten evaluated visual encoders, including settings in which REPA underperforms vanilla SiT. Combining REPI with REPA provides further gains for most encoders.

FID comparison across ten visual encoders, from DINOv2 through CLIP, WebSSL, I-JEPA, DeiT-III, MAE, and MoCoV3.
SiT-XL/2 · 100K steps · ImageNet 256 × 256 · without CFG.

Diffusion Backbones and Resolution

Results on DiT-L/2, DiG-L/2, and SiT-XL/2 at 512 resolution, from left to right.
From left to right: DiT-L/2, DiG-L/2, and SiT-XL/2 at 512 × 512. REPI consistently improves upon REPA across these architectures and settings.

Ablation Study

The temporary scaffold alone yields substantial improvements that persist after its removal. Adding the internalization objective further improves generation quality.

Component ablation: SiT-XL reaches FID 17.20, scaffold alone reaches 9.49, and scaffold plus internalization reaches 7.21 at 400K steps. With REPA, the corresponding values are 7.90, 6.54, and 6.33.
SiT-XL/2 · ImageNet 256 × 256 · without CFG.
Left: without REPA. Right: with REPA.

BibTeX

@misc{fu2026scaffold,
  title  = {Scaffold Then Internalize: Representation Injection for Diffusion Transformers},
  author = {Han Fu and Jiacheng Chen and Baoquan Zhao and Weidong Chen and Wei Liu and Qing Li and Xudong Mao},
  year   = {2026},
  eprint = {2609.35292},
  archivePrefix = {arXiv},
  primaryClass = {cs.CV},
  url    = {https://arxiv.org/abs/2609.35292}
}