AIAny
Icon for item

ALoDLM: Adaptively Looped Diffusion Language Models

Introduces a token-adaptive latent recurrence for diffusion language models that allocates computation per token during denoising. Key features: iterative latent refinement, discrete-feedback commits, and learned token-wise schedules to improve parallel decoding quality-efficiency trade-offs.

Introduction

Many diffusion-based language models (DLMs) promise much faster parallel decoding than autoregressive models, yet they often lag in quality. The surprising culprit is not model size but a computation–difficulty mismatch: some masked positions can be resolved quickly while others need substantially more iterative reasoning. ALoDLM flips the uniform compute assumption and adapts per-token computation during denoising, aiming to close the quality gap while preserving parallelism.

Key Findings
  • Token-adaptive latent recurrence: instead of applying the same denoising depth to all masked positions, ALoDLM repeatedly refines latent states for hard tokens while committing easy tokens early and feeding them back as discrete context. This reduces unnecessary computation on easy positions and concentrates iterative passes where they matter.
  • Joint learning of schedules: token-wise computation schedules are treated as latent variables and trained via a conditional negative evidence lower bound (NELBO), enabling the model to learn when to stop refining each token.
  • Strong empirical gains: trained at 1.7B and 8B scales, ALoDLM outperforms prior diffusion LMs and comparable autoregressive baselines on averaged metrics across eleven benchmarks, while retaining parallel decoding and yielding a favorable quality–efficiency trade-off under optimized inference.
  • Practical behavior: selective recurrence increases mask-to-mask interactions and stabilizes predictions at intermediate noise levels, allowing dynamic allocation of more loops where they yield the most benefit.
Who it's for and tradeoffs

Great fit if you are a researcher or engineer building parallel-generation LLMs who needs better generation quality without reverting to fully sequential decoding. ALoDLM is particularly relevant when inference engines can exploit parallelism and when per-step compute can be dynamically scheduled. Look elsewhere if you require the simplest implementation (ALoDLM adds scheduling and recurrence control logic) or if your deployment cannot support the optimized parallel inference needed to realize its efficiency gains.

Where it fits

ALoDLM sits between masked/diffusion LMs that apply uniform denoising and fully autoregressive decoders: it preserves much of the parallel throughput of diffusion approaches while closing the accuracy gap with AR methods. It complements other strategies such as looping early transformer layers (LoopMDM) or latent refinement frameworks, but distinctively learns token-wise stopping behavior.

How it works (brief)

At each denoising step ALoDLM runs iterative latent refinement passes. For each token position the model maintains a latent state and a learned decision mechanism: positions deemed ready are converted into discrete tokens and injected back as context; unresolved positions keep their latent states and get additional recurrent passes. Training formulates per-token compute schedules as latent variables and optimizes a conditional NELBO so token prediction and compute-allocation are learned jointly.

Practical considerations: implementing ALoDLM requires modifications to the denoising loop and a runtime that supports dynamic per-token recurrence and batched commits. The approach increases algorithmic complexity compared to fixed-depth DLMs but can yield substantial quality gains for parallel decoders when properly integrated with an optimized inference stack.

Information

  • Websitearxiv.org
  • AuthorsLiancheng Fang, Zhuowei Li, Youngeun Kim, Tianchen Zhao, Rajat Koner, Jiaye Wu, Linghan Xu, Xuanbai Chen, Xiang Xu, Zheng Zhang …
  • Published date2026/10/03

More Items

Adaptive-granularity post-retrieval evidence compressor for multimodal RAG that selects variable-sized regions from retrieved text, tables, images, and videos to retain necessary context while cutting irrelevant content. Key features include a hierarchy-based node encoder with parent-relative refinement and a critic that issues targeted follow-up retrievals; yields higher QA accuracy across five benchmarks and reduces reader-input tokens by ~14–28%.

Adapts LLM agents online by self-distilling verified execution trajectories into persistent LoRA weights during deployment to improve success and efficiency on long‑horizon tasks. Uses a frozen stable copy as a privileged teacher to predict hindsight next‑token distributions and filters invalid-action turns so experience consolidates without external solutions or memory retrieval.

Analyzes how LLM agents prefer items from particular sources during end-to-end search and how these preferences affect selections across shopping, accommodation, and scholarly domains. Shows source labels can override item quality and evaluates mitigation strategies such as hiding sources, relabeling, supplying missing information, and counter-prompts.