AIAny
Icon for item

SuperNav: An Agentic Navigation System for Any Task in Any Scene

Lets a pretrained multimodal LLM interpret navigation requests and orchestrate motion via tool calls for generalist robot navigation across unfamiliar scenes. Key features: an agent harness with Navigation Skills, a unified visual-point interface, task-progress tracking, and tool-based motion execution without navigation-specific fine-tuning.

Introduction

Most embodied navigation approaches fold motion behavior into fine-tuned vision-language models, making generalization rely heavily on navigation training data. SuperNav takes the opposite tack: keep the pretrained multimodal LLM focused on understanding requests, scene reasoning, and high-level decisions, and delegate physical motion to a tool-backed agent harness that the LLM controls.

Key Findings
  • Modular agent harness + tool calling: the system separates decision-making (MLLM) from motion execution (navigation tools). This means improvements to motion backends or recovery strategies can boost overall performance without retraining the LLM.
  • Unified visual-point interface: the model specifies destinations directly on observation images; the navigation tool executes movement and returns feedback. So what: the same decision interface works across different motion backends and reduces the need to predict low-level actions.
  • Strong empirical gains across tasks: SuperNav substantially outperforms evaluated baselines on instance-level, multi-object, and demand-driven navigation (e.g., 78.0% success on single-object navigation vs ~34% for top baselines) and shows competitive category-level results on HM3D without OVON-specific training. So what: the approach improves cross-scene and cross-task robustness rather than memorizing navigation behaviors.
Who it's for and tradeoffs

Great fit if you are building general-purpose service robots or embodied agents that must handle diverse human requests in novel environments and want a modular pipeline where language reasoning and motion control evolve independently. Look elsewhere if you need tightly optimized, end-to-end learned low-level control for a single, well-instrumented environment—SuperNav trades some per-domain optimality for broader generalization and modularity. Other practical tradeoffs: overall performance depends on the quality and latency of the motion backend and perception tools; deploying on physical robots requires calibration for sensing, actuation constraints, and safety checks.

How it works (brief)

The system equips an off-the-shelf MLLM with a harness that supplies: (1) agent-oriented Tools for observation, motion, and task management; (2) optional Navigation Skills that encode reusable procedural guidance (search, verification, recovery); and (3) task-progress and context tracking so the agent can sustain multi-step interactions. The visual-point interface grounds textual decisions in images, enabling closed-loop revisions from execution feedback.

Information

  • Websitearxiv.org
  • OrganizationsZhejiang University, Shenzhen University, Causa Robotics†Corresponding author
  • AuthorsJinkai Zhang, Jingyi Xu, Yuanhong Yu, Jiarui Guo, Ruizhen Hu, Hujun Bao, Xiaowei Zhou, Sida Peng
  • Published date2026/10/08

More Items

Creates programmable, real-time interactive code-based environments by separating deterministic simulator state from a shared neural video renderer so agents perceive, interact, and iteratively evolve via distilled playbooks; introduces Adversarial Forcing to distill a geometry-conditioned renderer for responsive visual feedback.

Measures how effectively LLM agents learn from interaction by playing 20 text-based games with novel or counterintuitive hidden rules, providing deterministic feedback, episode-wise scoring, and controlled variations to test retention and transfer.

Converts historical interaction traces into a reusable, queryable “worldbook” and runs a language-based world model agent (Trace2Env) as the environment for LLM agents — enabling stateful, grounded simulation with improved next-observation fidelity and long-horizon consistency.