Normalizes each domain's teacher feedback spread during multi-teacher on-policy distillation so that no domain (e.g., instruction following) overwhelms others, improving student recovery of specialist skills and raising average scores across benchmarks.
Improves test-time scaling of looped transformers by adaptively assigning extra recurrent iterations to tokens that benefit most. TaH2 is a post-training method that jointly trains an iteration decider with the backbone using lookahead depth supervision, boosting the accuracy–compute slope and peak accuracy on challenging benchmarks.