AIAny
Icon for item

PanoVLN: Towards Effective Panoramic Vision-and-Language Navigation

Uses 360° panoramic observations to improve vision-and-language navigation by predicting longer action sequences, confidence-guided execution, and combined semantic–geometric panorama features; yields large SR gains on R2R-CE and RxR-CE Val-Unseen.

Introduction

Wider visibility from a single 360° observation should reduce blind spots and simplify global route selection, yet naively swapping perspective images for panoramas provides only limited benefit. PanoVLN argues the gap comes from action prediction, supervision design, and visual representation — not just more pixels — and addresses each to unlock panoramas' potential for VLN.

Key Findings
  • Confidence-Guided Execution (CGE): the model predicts longer action sequences and dynamically decides how many predicted actions to execute before replanning, enabling larger turns and multi-step moves from a single panorama and improving planning stability.
  • Route-centric supervision: training routes are constructed with frequent branching points and clearer instructions so the model receives targeted supervision for route selection rather than only local grounding.
  • Semantic + geometric panorama features: combines RGB semantic cues with geometric layout signals from panoramas without expanding visual token counts, improving cross-direction spatial reasoning.
  • Strong empirical gains: using a 4B backbone and RGB-only input, PanoVLN improves success rate over prior state-of-the-art by 11.9% on R2R-CE Val-Unseen and 8.7% on RxR-CE Val-Unseen. Real-world tests on a quadruped robot show faster traversal and fewer pauses than prior VLN methods.
Who It's For and Trade-offs

Great fit if you research embodied navigation, multimodal VLN, or deploy agents with panoramic sensors and want methods that leverage full-surround context for longer-horizon planning. Look elsewhere if your platform cannot capture equirectangular panoramas, if compute/memory for longer action decoding is constrained, or if your use case strictly requires depth sensors (the main results use RGB-only fusion). The approach increases dataset and supervision complexity (branching routes, longer-horizon labels) and benefits environments where global layout cues matter more than fine local geometry.

Where It Fits

PanoVLN sits between classic perspective-view VLN and fully pano-native pipelines: it demonstrates that effective panoramic VLN requires adapting action execution and supervision, not merely feeding wider images into existing models. It is most relevant for embodied-AI research and robotic navigation settings where full-surround observations are available.

Information

  • Websitearxiv.org
  • OrganizationsZhejiang University, The University of Hong Kong
  • AuthorsZhen Wang, Changpeng Wang, Zhe Liu, Zhangyang Qi, Yuxiang Lu, Zimo Zeng, Donglian Qi, Xi Chen
  • Published date2026/09/28

More Items

Groups visually grounded appearances of the same physical instance into persistent, retrievable “biographies” so agents can follow objects across hours or days for long-video question answering. Links identity-aware observations to episodic context and visual evidence; improves EgoLifeQA to 72.0% and increases evidence-window reach from 37.6% to 58.9%.

Trains a single unified multimodal model with reinforcement learning to perform end-to-end self-reflection and iterative image repair — jointly learning the diagnostic (textual) reflection and the flow-based image revisions so credit propagates across rounds without an external verifier.

Uses diffusion-model forking moments as a proxy for perceptual distance to automatically generate pointwise reference-grounded labels, enabling annotation-free training of reference-based image quality assessment metrics.