AIAny
Icon for item

E-MoE: Enhanced Mixture-of-Experts for Non-Factorized Diffusion Language Models

Models sequence generation by unmasking multiple tokens per denoising step and replaces a factorized reverse process with a mixture over discrete routing-based latents from an MoE backbone; improves few-step sampling quality without increasing active parameters.

Introduction

Most masked diffusion sequence models treat the reverse (denoising) step as factorized across token positions, which hurts sample quality when using few denoising steps — the regime where diffusion gains over autoregressive decoding. E-MoE's core insight is to reuse the Mixture-of-Experts router as a discrete shared latent, turning the reverse process into a mixture of factorized distributions and thereby capturing cross-token correlations at no extra active-parameter cost.

Key Findings
  • Reusing MoE routing as the latent avoids adding a separate recognition model or designing a parametric prior, so training does not increase active parameters nor require auxiliary networks. This means practical scaling is aligned with existing MoE backbones.
  • Derives an ELBO for the mixture-based reverse process and a tractable latent-KL bound computable from the router itself, keeping optimization stable and interpretable during training.
  • Empirically improves few-step generation: on synthetic multi-modal benchmarks, binarized MNIST, and LM1B E-MoE markedly lowers generative perplexity at 1–2 denoising steps compared to factorized and continuous-latent baselines, and its latent remains active without warmup schedules or auxiliary losses.
  • The advantage decreases as the number of function evaluations (NFE) grows, consistent with the expectation that factorization becomes less limiting when each step unmasks fewer tokens.
Who it's for and tradeoffs

Great fit if you use or plan to use MoE-backed diffusion LMs and need better quality in low-NFE (few-step) parallel decoding. It offers improved sampling quality without extra active-parameter cost and avoids VAE-style posterior collapse. Look elsewhere if you cannot adopt an MoE backbone (method depends on routing as the latent), if very high-step diffusion is acceptable (advantage shrinks with many steps), or if your deployment strictly forbids sparse-expert architectures.

Method sketch

The method treats per-token, per-layer expert-routing choices as discrete latent codes. The model marginalizes over routes to couple tokens unmasked in the same step, overcoming the factorization error. Training optimizes a derived ELBO and a router-derived KL bound; at inference the same router provides the generative prior, so no extra recognition model or hand-designed prior is required.

Information

  • Websitearxiv.org
  • OrganizationsAffiliation: AXXX, Moscow, Russia, Affiliation: Applied AI Institute, Moscow, Russia
  • AuthorsArseny Ivanov, Alexander Kolesov, Alexander Korotin, Ivan Oseledets, Mikhail Goncharov
  • Published date2026/09/29

More Items

Calibrates multi-reward reinforcement learning by adaptively upweighting infrequently active rewards per rollout batch, so sparse objectives provide stronger signals when they matter. Proposes an inverse-square-root density correction and shows faster learning on tool-calling and math-reasoning tasks.

Demonstrates that pretrained transformers typically use only ~1–3 lines of depth to follow reference chains, and that a task‑trained rank‑8 LoRA applied at one early layer (with all other weights frozen) can extend reference‑following to dozens or hundreds of lines while adding only a few ten‑thousand parameters.

Investigates how rollout policy, token-level KL direction, and learning rate each affect LLM distillation across Llama3 and Qwen2.5 on reasoning tasks; finds KL direction and learning rate dominate outcomes while rollout policy has a modest effect.