AIAny
Icon for item

AgenticDataBench: A Comprehensive Benchmark for Data Agents

Provides a comprehensive benchmark to evaluate LLM-based data agents on realistic, multi-domain data-science workflows. Features skill-level ground-truth labels, 15 vertical domains (including real B2B tasks), and LLM-driven task generation to ensure coverage; includes an open testbed and agent evaluations.

Introduction

Automating end-to-end data-science workflows with LLM-based agents promises to cut manual effort, but progress is hampered by a lack of benchmarks that capture the diversity and skill composition of real tasks. AgenticDataBench aims to close that gap by assembling representative datasets, decomposing tasks into reusable data-science skills, and providing fine-grained labels so evaluations reflect practical agent capabilities rather than only coarse task success.

Key Findings
  • Multi-domain coverage: The benchmark collects tasks and datasets across 15 verticals, including five real-world B2B cases, so evaluations stress domain variance and business constraints rather than toy examples. This means agent results are more indicative of production readiness.
  • Skill-level labeling: Tasks are annotated by constituent data-science skills (extraction, cleaning, joining, modeling, interpretation), enabling per-skill performance analysis. This exposes which subroutines agents struggle with and guides targeted improvements.
  • Hybrid construction: For domains lacking real tasks, the authors use an LLM-based task-generation pipeline grounded in extracted skills to produce realistic workflows, expanding coverage without excessive manual effort. This trades perfect realism for scalability while retaining skill diversity.
  • Empirical evaluation: The paper evaluates state-of-the-art data agents on the benchmark and reports detailed, skill-wise failure modes rather than only aggregate scores, highlighting gaps in data integration, unstructured-text extraction, and multi-step reasoning.
Who it's for and tradeoffs

Great fit if you build or evaluate LLM-driven data agents, research agent architectures for data integration/analysis, or need a benchmark that surfaces per-skill weaknesses. Look elsewhere if you only need single-query SQL translation tests or tiny toy tables: AgenticDataBench emphasizes realistic, multi-step workflows and requires more complex testbeds and annotation effort. The LLM-generated tasks improve coverage but may not fully substitute high-fidelity proprietary business scenarios.

Information

  • Websitearxiv.org
  • AuthorsZhaoyan Sun, Shan Zhong, Daizhou Wen, Jiaxing Han, Guoliang Li, Ying Yan, Peng Zhang, Yu Su, Xiang Qi, Baolin Sun …
  • Published date2026/07/02

Categories

More Items

Provides a plug-and-play harness that makes existing agents omni-native by exposing hierarchical multimodal Skills, a standardized execution interface, dependency-aware orchestration, and a persistent Asset Registry. Represents multi-asset workflows as Declare Execution Graphs to schedule concurrent operations and enable cross-turn reuse across interchangeable execution backends.

Introduces an open-source multi-agent harness that automatically constructs, composes, and evolves modular model–harness units for long‑horizon, cross‑domain workflows. Key ingredients include a Host Agent for task decomposition and orchestration, EverOS for durable memory, and a Skill Forge of reusable procedures to improve task coverage via composition.

Learns joint predictive visual dynamics and action generation for generalist robot manipulation, translating future-relevant visual representations into actions. Integrates a Mixture-of-Transformers coupling a video expert and action expert, a frozen vision-language model for semantics, 4D distillation, and Causal Imprint; pretrained on a 20K+ hour heterogeneous corpus.