AIAny
Icon for item

FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders

Trains decoders and diffusion generators to be robust to random subsets of pretrained encoder layers by regularizing layer-fusion during training, narrowing the reconstruction–generation gap in representation autoencoders and improving ImageNet-256 PSNR and gFID.

Introduction

Most representation autoencoders (RAEs) reuse a frozen pretrained visual encoder but must choose which encoder layers to fuse into the latent. This fixed fusion creates a mismatch: shallow layers help pixel-accurate reconstruction while deep layers are easier for generative modeling. FuseReg’s core insight is simple and practical: train downstream models on randomized layer subsets so decoders and generators learn to be robust to cross-layer disagreement, narrowing the reconstruction–generation gap without changing the pretrained encoder.

Key Findings
  • Randomized layer-fusion regularization: training the decoder on normalized fusions from randomly sampled encoder-layer subsets explicitly penalizes sensitivity to cross-layer disagreement, enabling a single decoder to accept full, sparse, or single-layer fusions without retraining. So what: one downstream decoder becomes flexible across latent definitions used by generators.
  • Empirical gains on ImageNet-256: with DINOv3-L, a FuseReg decoder increases PSNR versus fixed-fusion decoders and, when replacing only the decoder, reduces unguided gFID by ~27% for k=23; joint regularization of decoder and DiT generator reduces unguided gFID by ~29% on DiT-Base. So what: both reconstruction quality and generation metrics improve simultaneously rather than trading off.
  • Theoretical justification: randomized fusion preserves the full-layer mean while converting cross-layer disagreement into an explicit regularization penalty and provides a second-order separation from deterministic global fusions. So what: the method is principled and not merely an empirical trick.
Who it's for and trade-offs

Great fit if you build image generation pipelines that reuse frozen encoders (RAEs/DiT) and you need a single downstream decoder that works across different fusion choices, or you want to boost generation metrics without re-training encoders. Look elsewhere if you can retrain or jointly finetune the encoder end-to-end for your task, or if your application strictly requires architecture-level fusion mechanisms (learnable gating) with different inductive biases; FuseReg modifies downstream training but does not change encoder weights.

Where it fits

FuseReg complements work on multi-layer fusion and richer visual tokenizers by addressing the downstream mismatch between reconstruction and generative modeling. It is applicable to RAE-based diffusion transformers and other pipelines that treat encoder features as latents, offering a low-cost modification (training-time sampling) with measurable gains in both reconstruction and generation metrics.

Information

  • Websitearxiv.org
  • OrganizationsUSC PSI Lab, Brown University, Rice University, University of Aberdeen, University of Notre Dame, University of Maryland, College Park, University of Pennsylvania
  • AuthorsHongyang Du, Yunfei Xie, Junjie Ye, Jiawei Yang, Xiaoyan Cong, Haodong Zhang, Yongchao Huang, Haiyu Wu, Zongxia Li, Shihang Gui …
  • Published date2026/09/25

More Items

Uses diffusion-model forking moments as a proxy for perceptual distance to automatically generate pointwise reference-grounded labels, enabling annotation-free training of reference-based image quality assessment metrics.

Learns joint predictive visual dynamics and action generation for generalist robot manipulation, translating future-relevant visual representations into actions. Integrates a Mixture-of-Transformers coupling a video expert and action expert, a frozen vision-language model for semantics, 4D distillation, and Causal Imprint; pretrained on a 20K+ hour heterogeneous corpus.

Introduces a rubric-based approach for video reward modeling that generates explicit, query-adaptive evaluation criteria before scoring to reduce scalar drift; proposes RGPO, a two-stage training (seed warm-up + joint optimization) that achieves strong pointwise and pairwise evaluation with high data efficiency.