AIAny
Icon for item

From Evidence to Action: How Tool-Using Agents Fail

Analyzes where LLM-based agents break the evidence-to-action chain and introduces SafeActBench, a 656-case, provenance-bound benchmark and deterministic evaluator to diagnose failures in investigation, timing, single-action execution, and multi-step workflows.

Introduction

Why this matters

Many deployed LLM agents do more than answer questions: they investigate, call tools, and change external state. This paper’s core insight is that correct end results are not enough — what matters for safety and correctness is whether the necessary evidence was established before each consequential action. The work shows that failures often originate upstream (incomplete investigation or premature action) rather than in the mechanics of single tool calls.

Key Findings
  • A diagnostic benchmark (SafeActBench) with 656 canonical cases across six operational domains reveals where the evidence-to-action chain breaks: static judgments can be strong while interactive execution is weak.
  • Evidence instrumentation matters: a provenance-bound Evidence Ledger plus a deterministic trajectory evaluator make it possible to verify what facts were established, when actions occurred, and whether downstream dependencies were satisfied.
  • Failure modes cluster upstream: many agents either stop investigation too early or act before required evidence exists; conditional on required evidence, single-action execution is generally reliable, but multi-action workflows expose unresolved prerequisites and incomplete propagation.
  • Evaluations across ten model–harness configurations (five model families paired with two harness styles) show substantial variation in investigation completeness and multi-step robustness, not just raw action correctness.
Who this is for and tradeoffs

Great fit if you need a repeatable, provenance-aware benchmark to evaluate or harden LLM-driven automation, especially for teams building transactional agents (customer ops, infra, healthcare, finance). It helps pinpoint whether failures are due to information gathering, evidence binding, timing, or workflow composition. Look elsewhere if you need deployment recipes, runtime tool libraries, or hands-on mitigation code — the paper diagnoses failures and provides evaluation infrastructure rather than turnkey runtime mitigations.

Where it sits in practice

This is primarily a research and evaluation contribution: use it to stress-test agents before granting them real-world effecting permissions. The combination of explicit evidence tracking and deterministic trajectory replay is the practical lever for turning opaque agent behavior into actionable diagnostics.

Information

  • Websitearxiv.org
  • OrganizationsPrinceton University, National University of Singapore, Hong Kong Baptist University, Amazon Web Services
  • AuthorsHongzhan Lin, Shidong Cao, Ziyang Luo, Wenhao Chai, Mong-Li Lee, Wynne Hsu
  • Published date2026/10/06

Categories

More Items

Verifies and preserves trajectory-derived skill edits for LLM agents by pairing each proposed edit with replayable execution evidence and re-executing the relevant trajectory segments. Introduces Replayable Evidence Cards, a replay-based verification gate, and a Provisional Edit Ledger to retain locally supported edits across epochs for continual skill evolution.

Analyzes how LLM agents prefer items from particular sources during end-to-end search and how these preferences affect selections across shopping, accommodation, and scholarly domains. Shows source labels can override item quality and evaluates mitigation strategies such as hiding sources, relabeling, supplying missing information, and counter-prompts.

Lets general-purpose vision-language models directly command robots via a compact mid-level action interface and asynchronous monitoring, enabling zero-shot manipulation without task-specific policy training; demonstrates strong sim benchmarks and real xArm6 transfer.