AIAny
Icon for item

MotorMind: Scaffolding General Vision Language Models for Zero-Shot Robot Manipulation

Lets general-purpose vision-language models directly command robots via a compact mid-level action interface and asynchronous monitoring, enabling zero-shot manipulation without task-specific policy training; demonstrates strong sim benchmarks and real xArm6 transfer.

Introduction

Why this matters

General-purpose vision-language models (VLMs) are rapidly improving, but most robotic manipulation systems either train specialized policies or wrap VLMs in heavy external tooling. MotorMind asks a different question: can an off-the-shelf VLM itself plan and adapt actions from observations when given an appropriate execution interface? The surprising answer is largely yes — with the right mid-level action abstraction and an asynchronous execution harness, a VLM can serve as the core decision-maker for zero-shot robot manipulation across simulation and a real robot.

Key Findings
  • Strong zero-shot performance: MotorMind reaches 66.7% success on the LIBERO-PRO base suite and 53.8% under task/object/position perturbations, versus prior zero-shot baselines at most ~13.3% and ~19.2% respectively. This shows a large absolute improvement from the interface and harness rather than model fine-tuning.
  • Real-world transfer: The same VLM-facing interface achieves 95% average success on a real xArm6 across direct manipulation and human-perturbation settings, indicating robust sim-to-real applicability without embodiment-specific policy training.
  • Component strengths and failure modes: Replacing the backbone VLM improves results; remaining failures are mainly visual grounding, embodied reasoning, and action-knowledge gaps. Local decisions show action selection is weaker than progress or completion judgments, identifying targets for future work.
  • Design principles: compact mid-level action representation (parameterized translations/rotations/gripper ops) + deterministic embodiment controllers, plus asynchronous background monitoring and memory summaries for non-blocking reconsideration.
Who it's for and tradeoffs

Great fit if you want to evaluate or deploy VLM-driven embodied decision-making without the cost of per-task policy learning, coding agents, or extensive grounding tools. The approach is useful for researchers benchmarking zero-shot manipulation, teams exploring rapid sim-to-real transfer, and practitioners seeking a modular VLM-to-robot interface.

Look elsewhere if your application requires millimeter-precise manipulation in visually ambiguous scenes today: MotorMind reduces the need for task-specific training but does not eliminate challenges in visual grounding, fine-grained contact reasoning, or rich motor primitives that specialist learned policies may handle better.

Short method note

MotorMind exposes a compact set of parameterized mid-level actions to the VLM; embodiment-specific controllers deterministically execute proposals. An asynchronous monitor checks updated observations and can cancel pending commands at action boundaries; memory summaries are generated in the background to inform subsequent planning. This keeps the VLM directly responsible for sequential decision-making while offloading low-level motion execution to deterministic controllers.

Information

  • Websitearxiv.org
  • OrganizationsUniversity of Illinois Urbana-Champaign
  • AuthorsBingxuan Li, Siqi Song, Yizhuo Wu, Jiarui Yao, Tong Zhang, Huan Zhang
  • Published date2026/09/29

More Items

Analyzes how LLM agents prefer items from particular sources during end-to-end search and how these preferences affect selections across shopping, accommodation, and scholarly domains. Shows source labels can override item quality and evaluates mitigation strategies such as hiding sources, relabeling, supplying missing information, and counter-prompts.

Predicts compact 'prospective tokens' that summarize upcoming information needs and uses them to select a small set of past frames for conditioning long-horizon video generation, improving long-range consistency, visual quality, and action alignment while remaining plug-and-play across diverse generators.

Turns live video into reusable textual memory and timely responses by training a streaming video LLM to proactively generate time-grounded captions and event summaries. Key components: Proactive Hierarchical Caption Memory (PHCM) for multi-scale records and Proactive State Transition Learning (PSTL) to balance response timing; trained on the OneStreamer-1M streaming dataset.