AIAny
Icon for item

Improving Test-Time Scaling with Adaptive Looped Transformers

Improves test-time scaling of looped transformers by adaptively assigning extra recurrent iterations to tokens that benefit most. TaH2 is a post-training method that jointly trains an iteration decider with the backbone using lookahead depth supervision, boosting the accuracy–compute slope and peak accuracy on challenging benchmarks.

Introduction

Most looped-transformer work evaluates fixed-depth recurrence, which wastes iterations on many tokens and obscures whether extra latent steps truly improve long-output inference. The core insight here is simple: letting the model focus extra iterations only on tokens where further recursion improves predictions yields both higher peak accuracy and a steeper accuracy-per-doubling-of-FLOPs slope.

Key Findings
  • TaH2: a lightweight post-training pipeline that jointly fine-tunes the backbone and an iteration decider using online "lookahead depth" supervision; this lets the model learn which tokens benefit from more loops.
  • Measured impact: on AIME benchmarks TaH2 raises the accuracy–compute slope by 53% (2.74 vs. 1.79) over the non-looped baseline and surpasses the baseline's peak accuracy by ~3.4 points at matched test-time compute.
  • Scalability with depth: while existing fixed-depth looped models largely plateau, TaH2's advantage grows as maximum iteration depth increases (gain vs. baseline: +2.8 points at depth 2 → +3.9 points at depth 8).
  • Practical trade: TaH2 improves attainable accuracy for test-time scalable inference without changing model size; it requires a small post-training stage and an inexpensive iteration-decider at inference.
Who it's for and trade-offs

Great fit if you need improved accuracy when scaling inference compute (longer outputs or optional extra FLOPs) but cannot increase model parameters. TaH2 is suited for teams that can afford a short post-training phase and want a predictable, token-adaptive compute/accuracy trade-off. Look elsewhere if you require a zero-change deployment (no extra decision logic) or absolute minimal inference-time latency overhead from any control logic—TaH2 adds a small runtime decision step and a bit of implementation complexity compared to fixed-depth looping.

Information

  • Websitearxiv.org
  • OrganizationsTsinghua University, Yale University
  • AuthorsYichen You, Tianyu Fu, Aosong Feng, Xingtai Lv, Xuefei Ning, Ning Ding, Yu Wang
  • Published date2026/09/28

More Items

Teaches small reasoning models to diagnose when additional internal thinking is insufficient and to selectively query stronger models; introduces FlyBy, a supervised + cost-aware RL framework that learns when/what to ask. Improves pass@ metrics on hard benchmarks while reducing serving cost.

Shows that private post-training changes leave measurable "behavioral shadows" in task-unrelated single-word outputs and proposes Active Taskless Distillation (ATD) to transfer capabilities to a student using only those single-word teacher responses, without teacher logits or parameters.

Normalizes each domain's teacher feedback spread during multi-teacher on-policy distillation so that no domain (e.g., instruction following) overwhelms others, improving student recovery of specialist skills and raising average scores across benchmarks.