Most representation autoencoders (RAEs) reuse a frozen pretrained visual encoder but must choose which encoder layers to fuse into the latent. This fixed fusion creates a mismatch: shallow layers help pixel-accurate reconstruction while deep layers are easier for generative modeling. FuseReg’s core insight is simple and practical: train downstream models on randomized layer subsets so decoders and generators learn to be robust to cross-layer disagreement, narrowing the reconstruction–generation gap without changing the pretrained encoder.
Key Findings
- Randomized layer-fusion regularization: training the decoder on normalized fusions from randomly sampled encoder-layer subsets explicitly penalizes sensitivity to cross-layer disagreement, enabling a single decoder to accept full, sparse, or single-layer fusions without retraining. So what: one downstream decoder becomes flexible across latent definitions used by generators.
- Empirical gains on ImageNet-256: with DINOv3-L, a FuseReg decoder increases PSNR versus fixed-fusion decoders and, when replacing only the decoder, reduces unguided gFID by ~27% for k=23; joint regularization of decoder and DiT generator reduces unguided gFID by ~29% on DiT-Base. So what: both reconstruction quality and generation metrics improve simultaneously rather than trading off.
- Theoretical justification: randomized fusion preserves the full-layer mean while converting cross-layer disagreement into an explicit regularization penalty and provides a second-order separation from deterministic global fusions. So what: the method is principled and not merely an empirical trick.
Who it's for and trade-offs
Great fit if you build image generation pipelines that reuse frozen encoders (RAEs/DiT) and you need a single downstream decoder that works across different fusion choices, or you want to boost generation metrics without re-training encoders. Look elsewhere if you can retrain or jointly finetune the encoder end-to-end for your task, or if your application strictly requires architecture-level fusion mechanisms (learnable gating) with different inductive biases; FuseReg modifies downstream training but does not change encoder weights.
Where it fits
FuseReg complements work on multi-layer fusion and richer visual tokenizers by addressing the downstream mismatch between reconstruction and generative modeling. It is applicable to RAE-based diffusion transformers and other pipelines that treat encoder features as latents, offering a low-cost modification (training-time sampling) with measurable gains in both reconstruction and generation metrics.