Many diffusion-based language models (DLMs) promise much faster parallel decoding than autoregressive models, yet they often lag in quality. The surprising culprit is not model size but a computation–difficulty mismatch: some masked positions can be resolved quickly while others need substantially more iterative reasoning. ALoDLM flips the uniform compute assumption and adapts per-token computation during denoising, aiming to close the quality gap while preserving parallelism.
Key Findings
- Token-adaptive latent recurrence: instead of applying the same denoising depth to all masked positions, ALoDLM repeatedly refines latent states for hard tokens while committing easy tokens early and feeding them back as discrete context. This reduces unnecessary computation on easy positions and concentrates iterative passes where they matter.
- Joint learning of schedules: token-wise computation schedules are treated as latent variables and trained via a conditional negative evidence lower bound (NELBO), enabling the model to learn when to stop refining each token.
- Strong empirical gains: trained at 1.7B and 8B scales, ALoDLM outperforms prior diffusion LMs and comparable autoregressive baselines on averaged metrics across eleven benchmarks, while retaining parallel decoding and yielding a favorable quality–efficiency trade-off under optimized inference.
- Practical behavior: selective recurrence increases mask-to-mask interactions and stabilizes predictions at intermediate noise levels, allowing dynamic allocation of more loops where they yield the most benefit.
Who it's for and tradeoffs
Great fit if you are a researcher or engineer building parallel-generation LLMs who needs better generation quality without reverting to fully sequential decoding. ALoDLM is particularly relevant when inference engines can exploit parallelism and when per-step compute can be dynamically scheduled. Look elsewhere if you require the simplest implementation (ALoDLM adds scheduling and recurrence control logic) or if your deployment cannot support the optimized parallel inference needed to realize its efficiency gains.
Where it fits
ALoDLM sits between masked/diffusion LMs that apply uniform denoising and fully autoregressive decoders: it preserves much of the parallel throughput of diffusion approaches while closing the accuracy gap with AR methods. It complements other strategies such as looping early transformer layers (LoopMDM) or latent refinement frameworks, but distinctively learns token-wise stopping behavior.
How it works (brief)
At each denoising step ALoDLM runs iterative latent refinement passes. For each token position the model maintains a latent state and a learned decision mechanism: positions deemed ready are converted into discrete tokens and injected back as context; unresolved positions keep their latent states and get additional recurrent passes. Training formulates per-token compute schedules as latent variables and optimizes a conditional NELBO so token prediction and compute-allocation are learned jointly.
Practical considerations: implementing ALoDLM requires modifications to the denoising loop and a runtime that supports dynamic per-token recurrence and batched commits. The approach increases algorithmic complexity compared to fixed-depth DLMs but can yield substantial quality gains for parallel decoders when properly integrated with an optimized inference stack.