AIAny
Icon for item

Raven: The Harness of Harnesses for Composable Agentic Intelligence

Introduces an open-source multi-agent harness that automatically constructs, composes, and evolves modular model–harness units for long‑horizon, cross‑domain workflows. Key ingredients include a Host Agent for task decomposition and orchestration, EverOS for durable memory, and a Skill Forge of reusable procedures to improve task coverage via composition.

Introduction

Most advances in single-agent systems don’t scale cleanly to long, cross‑domain workflows because harness complexity and domain coupling grow faster than manual design can manage. The paper's core insight is that treating each executable model+harness pair as a composable unit and giving the system mechanisms to assemble, archive, and evolve those units lets a host orchestrator expand reliable task coverage beyond individual agents under shared resource limits.

Key Findings
  • Composition as leverage: decomposing goals to specialized agents and composing their outputs increases reliable coverage on long‑horizon tasks, meaning tasks that none of the individual agents could handle end‑to‑end can succeed when orchestrated correctly. This shifts engineering effort from monolithic harness design to reusable harness components and wiring.
  • Persistent experience matters: EverOS provides durable user, agent, and world memory so successful workflows become reusable templates; making past experience queryable reduces repetitive design and speeds up future task assembly.
  • Automated evolution pipeline: the paper describes an Evolver that proposes, tests, and retains harness changes against benchmarks, enabling iterative harness improvement rather than one‑off manual tuning. This closes the loop from failure diagnosis to verified improvement.
Who it's for and tradeoffs

Great fit if you need to coordinate many specialized models or tools across multi‑step, long‑horizon workflows and want reusable automations rather than bespoke scripts. Expect extra system complexity: adopting the harness model requires building or integrating a catalog of specialized agents, defining evaluation benchmarks, and investing in durable memory and verification checks. Look elsewhere if your workload is single‑step or narrowly scoped—composition and evolution overhead may not pay off for trivial tasks.

Where it fits

This work positions itself between single, highly engineered agent systems and heavyweight workflow orchestration platforms: it is research‑oriented but also provided as an open implementation for teams that plan to operationalize multi‑agent orchestration with continuous improvement.

Information

  • Websitearxiv.org
  • Organizationshttps://evermind.ai/, https://github.com/EverMind-AI/Raven
  • AuthorsEverMind AI
  • Published date2026/09/27

Categories

More Items

Provides a plug-and-play harness that makes existing agents omni-native by exposing hierarchical multimodal Skills, a standardized execution interface, dependency-aware orchestration, and a persistent Asset Registry. Represents multi-asset workflows as Declare Execution Graphs to schedule concurrent operations and enable cross-turn reuse across interchangeable execution backends.

Learns joint predictive visual dynamics and action generation for generalist robot manipulation, translating future-relevant visual representations into actions. Integrates a Mixture-of-Transformers coupling a video expert and action expert, a frozen vision-language model for semantics, 4D distillation, and Causal Imprint; pretrained on a 20K+ hour heterogeneous corpus.

Alternates a Planner (issues sub-queries) and a Synthesizer (integrates retrieved evidence into a persistent summary) to tackle long-horizon deep-search; introduces Role‑Decoupled Policy Optimization (RDPO) for role-specific RL credit assignment and shows strong results (IterSynth-8B reaches 50.7% on five benchmarks).