AIAny
Icon for item

Agent Priors-guided Policy Learning

Records structural priors with skill-specific policies so a runtime agent can select and compose the version of each skill best suited to new states, improving out-of-distribution and compositional generalization for robot manipulation from few demonstrations.

Introduction

Robots that learn from only a few demonstrations must generalize both how individual skills behave in new situations and how skills can be recombined. The core insight of this paper is to keep each policy's training-time structural assumption — a "structural prior" specifying what the behavior depends on — as part of the skill's interface so a task-level caller can choose the right variant when composing skills.

Key Findings
  • Introducing APPL (Agent Priors-guided Policy Learning), which trains multiple prior-specific versions per skill and records each version's prior, handoffs, training support, and simple verification evidence. A separate runtime agent then selects among these frozen policies and composes them toward new goals.
  • In adapted MetaWorld tasks with very few demonstrations, the best APPL candidate achieved 89.58% macro out-of-distribution success versus 37.92% for a fixed relational-prior baseline and 28.96% for a vanilla diffusion policy—showing large gains in spatial/extrapolation generalization.
  • On long-horizon ManiSkill scenarios, APPL scored 50.0% on motion-OOD cases, 92.5% on task-level variants, and solved 8 of 16 novel compositions. Ablations that hid interface information or priors substantially reduced motion-, task-level, and composition performance.
  • The paper emphasizes that verification evidence is limited (verification uses demonstrated entry states and short rollouts) and that stated applicability is an hypothesis rather than a guarantee of competence.
Who it's for and trade-offs

Great fit if you research or build robot systems that must reuse a small set of demonstrated multi-step behaviors across new spatial layouts or reordered sequences; APPL is specifically targeted at few-shot demonstration regimes where separating skill learning and composition helps. Look elsewhere if you need formal guarantees of safety or recovery behaviors out-of-domain, if you require full end-to-end training without modularization, or if your environment cannot be simulated for verification—APPL's results are presented in simulated manipulation benchmarks and its verification is heuristic.

Method overview

A construction agent segments demonstrations into overlapping skill segments, proposes multiple structural priors per segment, trains one policy per prior, and records verification evidence. At runtime a separate agent reads those records, selects a prior-specific policy with its arguments and stop conditions, and composes policies to achieve new task goals. The approach trades the simplicity of single-policy interfaces for richer metadata that guides selection and composition under distribution shifts.

Information

  • Websitearxiv.org
  • OrganizationsAffiliation: Zhiwei Xue, Affiliation: Xinhu Li, Affiliation: Harold Soh, Affiliation: National University of Singapore, Affiliation: [email protected], Affiliation: [email protected]
  • AuthorsPuming Jiang, Tianrun Hu, Haozhe Du, Yibo Li, Zhiwei Xue, Xinhu Li, Harold Soh
  • Published date2026/09/28

More Items

Constructs and continually maintains explicit belief states for long-horizon LLM agents, combining a structured world estimate with unresolved epistemic and achievement gaps. Adds consistency validation, Belief Trapping detection, and tailored recovery to improve execution and diagnosis benchmarks.

Analyzes how proposer–solver loops in self-evolving search agents can develop shared errors (co-cheating) that inflate internal rewards; introduces Multi-Sample Verification and CrossFit (cross-fitted scoring with partitioned sources) to reduce false agreement and improve downstream search performance.

Co-evolves candidate solutions and web-search queries to help LLM-driven evolutionary discovery, using a retrieval gate plus bilevel inner/outer loops that refine queries, rank documents by predicted solution value, and generate evaluated candidates.