Most deployed LLM agents face steady streams of related tasks and only receive a sparse verification signal at episode end; yet their own execution traces are an underused source of supervision. ASCENT proposes a single-pass, online internalization method that turns these verified trajectories into persistent model improvements without requiring external solutions or explicit memory retrieval.
Key Findings
- Self-distillation via a frozen stable copy: a frozen initial model receives the executed trajectory plus its verification outcome as privileged hindsight and predicts next-token distributions along that trajectory. Distilling those distributions into LoRA fast weights updates the live agent for later tasks, so verified outcomes teach the policy without imitating potentially unstable raw generations.
- Invalid-action removal improves signal quality: filtering out turns that correspond to invalid or impossible actions yields stronger privileged targets and speeds up consolidation of useful behaviors.
- Works in a single pass over a stream of tasks: the method updates weights online as experience accumulates, improving task success and interaction efficiency across embodied and web-agent benchmarks (ALFWorld, WebShop, AppWorld) and transferring to held-out scenes.
- Robust to sparse outcome supervision but has limits: ASCENT characterizes the population target under sparse verification and shows gains when outcome signals are informative, while identifying regimes where extremely sparse or noisy verification constrains learning.
Who it's for and trade-offs
Great fit if you run LLM-based agents that interact over long episodes and can obtain a sparse verification signal at episode end — ASCENT lets you consolidate verified execution experience into model parameters during deployment without storing and retrieving long textual memories. Look elsewhere if you cannot provide any reliable outcome/verification signal, if you must strictly avoid on-line weight updates for compliance reasons, or if you need guarantees against any catastrophic policy drift from in‑deployment updates (ASCENT mitigates but does not eliminate such risks).