Most masked diffusion sequence models treat the reverse (denoising) step as factorized across token positions, which hurts sample quality when using few denoising steps — the regime where diffusion gains over autoregressive decoding. E-MoE's core insight is to reuse the Mixture-of-Experts router as a discrete shared latent, turning the reverse process into a mixture of factorized distributions and thereby capturing cross-token correlations at no extra active-parameter cost.
Key Findings
- Reusing MoE routing as the latent avoids adding a separate recognition model or designing a parametric prior, so training does not increase active parameters nor require auxiliary networks. This means practical scaling is aligned with existing MoE backbones.
- Derives an ELBO for the mixture-based reverse process and a tractable latent-KL bound computable from the router itself, keeping optimization stable and interpretable during training.
- Empirically improves few-step generation: on synthetic multi-modal benchmarks, binarized MNIST, and LM1B E-MoE markedly lowers generative perplexity at 1–2 denoising steps compared to factorized and continuous-latent baselines, and its latent remains active without warmup schedules or auxiliary losses.
- The advantage decreases as the number of function evaluations (NFE) grows, consistent with the expectation that factorization becomes less limiting when each step unmasks fewer tokens.
Who it's for and tradeoffs
Great fit if you use or plan to use MoE-backed diffusion LMs and need better quality in low-NFE (few-step) parallel decoding. It offers improved sampling quality without extra active-parameter cost and avoids VAE-style posterior collapse. Look elsewhere if you cannot adopt an MoE backbone (method depends on routing as the latent), if very high-step diffusion is acceptable (advantage shrinks with many steps), or if your deployment strictly forbids sparse-expert architectures.
Method sketch
The method treats per-token, per-layer expert-routing choices as discrete latent codes. The model marginalizes over routes to couple tokens unmasked in the same step, overcoming the factorization error. Training optimizes a derived ELBO and a router-derived KL bound; at inference the same router provides the generative prior, so no extra recognition model or hand-designed prior is required.