Robots that learn from only a few demonstrations must generalize both how individual skills behave in new situations and how skills can be recombined. The core insight of this paper is to keep each policy's training-time structural assumption — a "structural prior" specifying what the behavior depends on — as part of the skill's interface so a task-level caller can choose the right variant when composing skills.
Key Findings
- Introducing APPL (Agent Priors-guided Policy Learning), which trains multiple prior-specific versions per skill and records each version's prior, handoffs, training support, and simple verification evidence. A separate runtime agent then selects among these frozen policies and composes them toward new goals.
- In adapted MetaWorld tasks with very few demonstrations, the best APPL candidate achieved 89.58% macro out-of-distribution success versus 37.92% for a fixed relational-prior baseline and 28.96% for a vanilla diffusion policy—showing large gains in spatial/extrapolation generalization.
- On long-horizon ManiSkill scenarios, APPL scored 50.0% on motion-OOD cases, 92.5% on task-level variants, and solved 8 of 16 novel compositions. Ablations that hid interface information or priors substantially reduced motion-, task-level, and composition performance.
- The paper emphasizes that verification evidence is limited (verification uses demonstrated entry states and short rollouts) and that stated applicability is an hypothesis rather than a guarantee of competence.
Who it's for and trade-offs
Great fit if you research or build robot systems that must reuse a small set of demonstrated multi-step behaviors across new spatial layouts or reordered sequences; APPL is specifically targeted at few-shot demonstration regimes where separating skill learning and composition helps. Look elsewhere if you need formal guarantees of safety or recovery behaviors out-of-domain, if you require full end-to-end training without modularization, or if your environment cannot be simulated for verification—APPL's results are presented in simulated manipulation benchmarks and its verification is heuristic.
Method overview
A construction agent segments demonstrations into overlapping skill segments, proposes multiple structural priors per segment, trains one policy per prior, and records verification evidence. At runtime a separate agent reads those records, selects a prior-specific policy with its arguments and stop conditions, and composes policies to achieve new task goals. The approach trades the simplicity of single-policy interfaces for richer metadata that guides selection and composition under distribution shifts.