AIAny
Icon for item

Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It

Demonstrates that pretrained transformers typically use only ~1–3 lines of depth to follow reference chains, and that a task‑trained rank‑8 LoRA applied at one early layer (with all other weights frozen) can extend reference‑following to dozens or hundreds of lines while adding only a few ten‑thousand parameters.

Introduction

Oops! Something went wrong

[next-mdx-remote-client] error compiling MDX: Unexpected character `<` (U+003C) before name, expected a character that can start a name, such as a letter, `$`, or `_` More information: https://mdxjs.com/docs/troubleshooting-mdx

Information

  • Websitearxiv.org
  • OrganizationsGeorgia Institute of Technology
  • AuthorsZehao Jin, Ruixuan Deng, Junran Wang
  • Published date2026/09/29

More Items

Calibrates multi-reward reinforcement learning by adaptively upweighting infrequently active rewards per rollout batch, so sparse objectives provide stronger signals when they matter. Proposes an inverse-square-root density correction and shows faster learning on tool-calling and math-reasoning tasks.

Models sequence generation by unmasking multiple tokens per denoising step and replaces a factorized reverse process with a mixture over discrete routing-based latents from an MoE backbone; improves few-step sampling quality without increasing active parameters.

Investigates how rollout policy, token-level KL direction, and learning rate each affect LLM distillation across Llama3 and Qwen2.5 on reasoning tasks; finds KL direction and learning rate dominate outcomes while rollout policy has a modest effect.