Why this matters
Many deployed LLM agents do more than answer questions: they investigate, call tools, and change external state. This paper’s core insight is that correct end results are not enough — what matters for safety and correctness is whether the necessary evidence was established before each consequential action. The work shows that failures often originate upstream (incomplete investigation or premature action) rather than in the mechanics of single tool calls.
Key Findings
- A diagnostic benchmark (SafeActBench) with 656 canonical cases across six operational domains reveals where the evidence-to-action chain breaks: static judgments can be strong while interactive execution is weak.
- Evidence instrumentation matters: a provenance-bound Evidence Ledger plus a deterministic trajectory evaluator make it possible to verify what facts were established, when actions occurred, and whether downstream dependencies were satisfied.
- Failure modes cluster upstream: many agents either stop investigation too early or act before required evidence exists; conditional on required evidence, single-action execution is generally reliable, but multi-action workflows expose unresolved prerequisites and incomplete propagation.
- Evaluations across ten model–harness configurations (five model families paired with two harness styles) show substantial variation in investigation completeness and multi-step robustness, not just raw action correctness.
Who this is for and tradeoffs
Great fit if you need a repeatable, provenance-aware benchmark to evaluate or harden LLM-driven automation, especially for teams building transactional agents (customer ops, infra, healthcare, finance). It helps pinpoint whether failures are due to information gathering, evidence binding, timing, or workflow composition. Look elsewhere if you need deployment recipes, runtime tool libraries, or hands-on mitigation code — the paper diagnoses failures and provides evaluation infrastructure rather than turnkey runtime mitigations.
Where it sits in practice
This is primarily a research and evaluation contribution: use it to stress-test agents before granting them real-world effecting permissions. The combination of explicit evidence tracking and deterministic trajectory replay is the practical lever for turning opaque agent behavior into actionable diagnostics.