AIAny
Icon for item

Omni-IO Skills: Harnessing Your Agent Omni-Native

Provides a plug-and-play harness that makes existing agents omni-native by exposing hierarchical multimodal Skills, a standardized execution interface, dependency-aware orchestration, and a persistent Asset Registry. Represents multi-asset workflows as Declare Execution Graphs to schedule concurrent operations and enable cross-turn reuse across interchangeable execution backends.

Introduction

Most agent research ties new modalities to expensive model updates or brittle ad-hoc tool assemblies. Omni-IO Skills argues an alternative: compose capability at the harness level so a host agent gains broad, evolvable multimodal production abilities without changing its reasoning core.

Key Findings
  • Harness design: 27 hierarchical Skills cover 38 representative tasks across seven artifact modalities (text, image, audio, video, documents, 3D assets, code) and four capability families (understanding, generation, reasoning, retrieval). This modularity lets different execution backends be swapped without changing the host agent.
  • Execution model: Workflows are encoded as Declare Execution Graphs, enabling dependency-aware scheduling that runs independent operations in parallel and registers successful outputs to a persistent Asset Registry for downstream and cross-turn reuse.
  • Empirical gains: On the UniM-90 benchmark, the harness raised input-support rates of two host agents to 100% and substantially improved semantic-quality and strict-structure scores, demonstrating that harness-level composition can yield production-grade multimodal behavior without retraining the core model.
Who it's for and trade-offs

Great fit if you want to extend an existing LLM-based agent to handle diverse media and multi-step asset workflows quickly, or if you need replaceable execution backends and reusable artifacts across turns. Look elsewhere if your priority is improving the model’s internal multimodal reasoning via end-to-end training, or if you cannot accept the added engineering surface (runtime orchestration, asset registry maintenance, and backend adapters) that a harness introduces.

Information

  • Websitearxiv.org
  • OrganizationsNational University of Singapore, University of Oxford
  • AuthorsYanlin Li, Mingyang Hao, Shengqiong Wu, Hao Fei, Mong-Li Lee, Wynne Hsu
  • Published date2026/09/25

Categories

More Items

Introduces an open-source multi-agent harness that automatically constructs, composes, and evolves modular model–harness units for long‑horizon, cross‑domain workflows. Key ingredients include a Host Agent for task decomposition and orchestration, EverOS for durable memory, and a Skill Forge of reusable procedures to improve task coverage via composition.

Learns joint predictive visual dynamics and action generation for generalist robot manipulation, translating future-relevant visual representations into actions. Integrates a Mixture-of-Transformers coupling a video expert and action expert, a frozen vision-language model for semantics, 4D distillation, and Causal Imprint; pretrained on a 20K+ hour heterogeneous corpus.

Alternates a Planner (issues sub-queries) and a Synthesizer (integrates retrieved evidence into a persistent summary) to tackle long-horizon deep-search; introduces Role‑Decoupled Policy Optimization (RDPO) for role-specific RL credit assignment and shows strong results (IterSynth-8B reaches 50.7% on five benchmarks).