Papers
arxiv:2609.31620

FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders

Published on Sep 25
· Submitted by
Hongyang Du
on Sep 28
#1 Paper of the day
Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,

Abstract

Representation autoencoders (RAEs) reuse features from a pretrained visual encoder as reconstruction and diffusion latents, integrating strong visual representations into image generation. However, RAEs still need to decide which encoder layers form the shared latent space for the generator and pixel decoder. This choice involves a trade-off. Shallower layers tend to preserve fine pixel details better, while deeper layers tend to yield better generation metrics. A fixed heuristic layer fusion therefore couples two stages that benefit from different information. We introduce FuseReg, which replaces heuristic feature selection with training over random subsets of encoder layers. We theoretically analyze the underlying mechanism: subset sampling explicitly penalizes sensitivity to cross-layer disagreement. On ImageNet-256 with DINOv3-L, a single FuseReg decoder reconstructs from full, sparse, and single-layer fusions without retraining, achieving higher PSNR than decoders specialized to fixed fusions. This flexibility also benefits generation: decoder replacement alone reduces unguided gFID by 27% with an unchanged RAEv2 DiT-XL generator. The same regularization principle extends to diffusion training, with joint regularization of both stages reducing unguided gFID by 29% on DiT-Base. These results show that training downstream models for layer-fusion robustness narrows the reconstruction-generation gap without modifying the pretrained encoder.

Community

Paper author Paper submitter

FuseReg: Mitigating the Reconstruction–Generation Gap in RAEs

Representation Autoencoders fuse features from multiple encoder layers into a shared latent space.

However, reconstruction and generation prefer different parts of the hierarchy:

  • Decoder: prefers shallow, pixel-rich features.
  • DiT: prefers deeper, more structured semantic features.

This creates a reconstruction–generation mismatch.

FuseReg

FuseReg replaces fixed layer fusion with random layer-subset sampling during training.

For a sampled subset (S),

zS=1∣S∣∑k∈Shk. z_S = \frac{1}{|S|}\sum_{k\in S} h_k.

The normalized sampling keeps the latent mean unchanged:

E[zS∣x]=hˉ. \mathbb{E}[z_S \mid x] = \bar h.

Under a linear squared-loss surrogate, we further show that random subset sampling explicitly regularizes cross-layer disagreement.

In short, FuseReg discourages models from relying too strongly on any particular layer composition.


Image_20260927144217_3069_29

Robust across layer fusions

A single FuseReg decoder remains effective across full, sparse, and even single-layer inputs, while fixed-fusion decoders degrade strongly outside their training fusion.


Image_20260927144239_3070_29

Generation improvements

FuseReg can be applied independently to the decoder and DiT.

  • Decoder replacement only: unguided gFID 3.01 → 2.21
  • DiT-Base, joint FuseReg: unguided gFID 13.96 → 9.93
  • DiT-XL: unguided gFID 2.91 → 2.38

The same trend also holds across different pretrained encoder families.


Image_20260927144257_3071_29

One decoder, many layer combinations

The same FuseReg decoder can reconstruct from different layer subsets without retraining, including sparse and single-layer inputs.


Image_20260927144317_3072_29

What changes?

FuseReg distributes reconstruction information more evenly across encoder depth:

  • lower dependence on individual layers;
  • stronger reconstruction from single layers;
  • more stable representations under changing layer compositions.

Instead of searching for one optimal fusion, FuseReg trains the model to remain robust across a distribution of layer fusions.

This is an automated message from the ResearchStudio team.

We created an interactive ResearchStudio Reel for this paper. It includes a visual poster, a video, and a blog, all available for download in editable formats.

Visual poster for this paper

Open the ResearchStudio Reel →

Download all files from Hugging Face

Please give this comment a thumbs up if you find the Reel helpful!

Want to explore or create Reels for more papers? Visit the ResearchStudio demo.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.31620
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 1

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.31620 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.31620 in a Space README.md to link it from this page.

Collections including this paper 2