Verifies and preserves trajectory-derived skill edits for LLM agents by pairing each proposed edit with replayable execution evidence and re-executing the relevant trajectory segments. Introduces Replayable Evidence Cards, a replay-based verification gate, and a Provisional Edit Ledger to retain locally supported edits across epochs for continual skill evolution.
Analyzes where LLM-based agents break the evidence-to-action chain and introduces SafeActBench, a 656-case, provenance-bound benchmark and deterministic evaluator to diagnose failures in investigation, timing, single-action execution, and multi-step workflows.