Most looped-transformer work evaluates fixed-depth recurrence, which wastes iterations on many tokens and obscures whether extra latent steps truly improve long-output inference. The core insight here is simple: letting the model focus extra iterations only on tokens where further recursion improves predictions yields both higher peak accuracy and a steeper accuracy-per-doubling-of-FLOPs slope.
Key Findings
- TaH2: a lightweight post-training pipeline that jointly fine-tunes the backbone and an iteration decider using online "lookahead depth" supervision; this lets the model learn which tokens benefit from more loops.
- Measured impact: on AIME benchmarks TaH2 raises the accuracy–compute slope by 53% (2.74 vs. 1.79) over the non-looped baseline and surpasses the baseline's peak accuracy by ~3.4 points at matched test-time compute.
- Scalability with depth: while existing fixed-depth looped models largely plateau, TaH2's advantage grows as maximum iteration depth increases (gain vs. baseline: +2.8 points at depth 2 → +3.9 points at depth 8).
- Practical trade: TaH2 improves attainable accuracy for test-time scalable inference without changing model size; it requires a small post-training stage and an inexpensive iteration-decider at inference.
Who it's for and trade-offs
Great fit if you need improved accuracy when scaling inference compute (longer outputs or optional extra FLOPs) but cannot increase model parameters. TaH2 is suited for teams that can afford a short post-training phase and want a predictable, token-adaptive compute/accuracy trade-off. Look elsewhere if you require a zero-change deployment (no extra decision logic) or absolute minimal inference-time latency overhead from any control logic—TaH2 adds a small runtime decision step and a bit of implementation complexity compared to fixed-depth looping.