Most experience-driven attempts to evolve agent skills treat an edited revision as an all-or-nothing unit: if the revision fails global validation its constituent edits and their behavioral evidence are discarded. EVISKILL challenges that assumption by insisting each proposed edit be grounded in the execution context that motivated it, re-executed under the edited skill, and tracked across epochs so locally supported corrections can survive and be refined before global commitment. The core insight is that behavioral verification plus selective persistence of partial progress yields more reliable skill accumulation than blind textual edits or single-shot revision gates.
Key Findings
- Empirical wins across interactive benchmarks: EVISKILL reaches top accuracy in the majority of model–dataset settings (best in 14 of 18), improving over no-skill baselines by an average of ~17.9 percentage points — this shows replay-grounded edits materially improve task success.
- Replay filters and repairs common false-positive edits: on some environments 20–42% of candidate edits fail initial replay (i.e., they look plausible but do not change execution); targeted replay plus reflect-and-repair recovers many (e.g., 49/54 reflected edits repaired in a reported case).
- Provisioning preserves useful progress: when a revision is globally rejected, locally supported edits retained in the Provisional Edit Ledger later enter the validated skill at high rates (reported ~72.9% of provisionally retained edits succeed eventually). Cross-epoch evidence reuse is nontrivial (about 23% of unique evidence cards reused across epochs).
Who it's for and tradeoffs
Great fit if you build or evaluate LLM agents that learn procedural/interaction skills from trajectories and need edits that are behaviorally validated rather than just textually plausible. It helps research and development focused on continual skill accumulation, reproducible edit auditing, and robust evaluation on interactive benchmarks. Look elsewhere if you need parameter-level fine-tuning, very low-latency online edits (replay introduces extra verification cost), or when your task lacks meaningful short replay segments to validate edits.
Mechanisms (brief)
- Replayable Evidence Cards pair a proposed correction with the minimal trigger ranges and execution context needed to reconstruct the state for re-execution; replay executes only short segments to check behavioral change.
- A replay-based gate returns accept/reflect/reject decisions; reflected edits can be repaired via feedback. Accepted edits are recorded in a Provisional Edit Ledger so they persist across epochs even if a sibling revision is later rejected. The agent executes a Working Skill composed of the Validated Skill plus the Provisional Ledger until global validation commits updates.