AIAny
Icon for item

ASCENT: Online Test-Time Training of Long-Horizon Agents via Self-Distillation of Verified Experience

Adapts LLM agents online by self-distilling verified execution trajectories into persistent LoRA weights during deployment to improve success and efficiency on long‑horizon tasks. Uses a frozen stable copy as a privileged teacher to predict hindsight next‑token distributions and filters invalid-action turns so experience consolidates without external solutions or memory retrieval.

Introduction

Most deployed LLM agents face steady streams of related tasks and only receive a sparse verification signal at episode end; yet their own execution traces are an underused source of supervision. ASCENT proposes a single-pass, online internalization method that turns these verified trajectories into persistent model improvements without requiring external solutions or explicit memory retrieval.

Key Findings
  • Self-distillation via a frozen stable copy: a frozen initial model receives the executed trajectory plus its verification outcome as privileged hindsight and predicts next-token distributions along that trajectory. Distilling those distributions into LoRA fast weights updates the live agent for later tasks, so verified outcomes teach the policy without imitating potentially unstable raw generations.
  • Invalid-action removal improves signal quality: filtering out turns that correspond to invalid or impossible actions yields stronger privileged targets and speeds up consolidation of useful behaviors.
  • Works in a single pass over a stream of tasks: the method updates weights online as experience accumulates, improving task success and interaction efficiency across embodied and web-agent benchmarks (ALFWorld, WebShop, AppWorld) and transferring to held-out scenes.
  • Robust to sparse outcome supervision but has limits: ASCENT characterizes the population target under sparse verification and shows gains when outcome signals are informative, while identifying regimes where extremely sparse or noisy verification constrains learning.
Who it's for and trade-offs

Great fit if you run LLM-based agents that interact over long episodes and can obtain a sparse verification signal at episode end — ASCENT lets you consolidate verified execution experience into model parameters during deployment without storing and retrieving long textual memories. Look elsewhere if you cannot provide any reliable outcome/verification signal, if you must strictly avoid on-line weight updates for compliance reasons, or if you need guarantees against any catastrophic policy drift from in‑deployment updates (ASCENT mitigates but does not eliminate such risks).

Information

  • Websitearxiv.org
  • OrganizationsUniversity of New South Wales (UNSW Sydney)
  • AuthorsHaodong Lu, Dong Gong
  • Published date2026/10/04

More Items

Introduces a token-adaptive latent recurrence for diffusion language models that allocates computation per token during denoising. Key features: iterative latent refinement, discrete-feedback commits, and learned token-wise schedules to improve parallel decoding quality-efficiency trade-offs.

Turns solved protein structures into FoldingCorpus and Fold2Reason — a post-training recipe that supervises an LLM with discrete structural Q&A plus continuous 3D geometry to improve spatial, graph and scientific reasoning; reports consistent gains across 10 benchmarks.

Calibrates multi-reward reinforcement learning by adaptively upweighting infrequently active rewards per rollout batch, so sparse objectives provide stronger signals when they matter. Proposes an inverse-square-root density correction and shows faster learning on tool-calling and math-reasoning tasks.