Intelligent Visual Perception

Zero-Shot Target Navigation Integrating Visual Scene Understanding and Commonsense Reasoning

  • ZHOU Feiyu ,
  • ZHANG Xing
Expand
  • 1. School of Architecture & Urban Planning, Shenzhen University, Shenzhen 518060, Guangdong, China;
    2. Guangdong Key Laboratory of Urban Informatics, Shenzhen University, Shenzhen 518060, Guangdong, China;
    3. Key Laboratory for Geo-Environmental Monitoring of Great Bay Area, MNR, Shenzhen University, Shenzhen 518060, Guangdong, China

Received date: 2026-01-30

  Online published: 2026-08-01

Abstract

To address the problem that existing zero-shot navigation methods are difficult to simultaneously satisfy complex instruction understanding and efficient visual search, a zero-shot target navigation framework integrating the dual-level commonsense reasoning of large language models and the synergistic environmental perception of vision-language models was proposed. Through a modular design, the large language model was positioned as a commonsense reasoning engine to parse the target from natural language instructions and generate dual-level “room-object” spatial search priors. The vision-language model was taken as a real-time scene interpreter, which performed semantic matching based on this prior and output local semantic values. The two models collaborated to construct a semantic value map, achieving an organic unification of the commonsense-guided global search and the vision-driven local localization. Experimental results indicate that the proposed method improves the navigation success rate and the path efficiency. Its outstanding performance on navigation efficiency metrics validates the effectiveness of the spatial prior reasoning of large language models in improving search efficiency. Under the zero-shot setting, this framework achieves a deep integration of real-time perception and commonsense generalization, providing an efficient, accurate, and practical solution for intelligent navigation in open scenarios.

Cite this article

ZHOU Feiyu , ZHANG Xing . Zero-Shot Target Navigation Integrating Visual Scene Understanding and Commonsense Reasoning[J]. Journal of Applied Sciences, 2026 , 44(4) : 615 -625 . DOI: 10.3969/j.issn.0255-8297.2026.04.007

References

[1] Batra D, Gokaslan A, Kembhavi A, et al. Objectnav revisited:on evaluation of embodied agents navigating to objects[PP/OL]. V2(2020-08-30)[2026-01-30].https://doi.org/10.48550/arXiv.2006.13171.
[2] 谢远龙,王书亭,程祥,等.大模型驱动的具身智能机器人导航技术综述[J].华中师范大学学报(自然科学版), 2025, 59(5):677-693.Xie Y L, Wang S T, Cheng X, et al. An overview of large model-driven embodied intelligent navigation[J]. Journal of Central China Normal University(Natural Sciences), 2025, 59(5):677-693.(in Chinese)
[3] 高宇宁,王安成,赵华凯,等.基于深度强化学习的视觉导航方法综述[J].计算机工程与应用, 2025,61(10):66-78.Gao Y N, Wang A C, Zhao H K, et al. Review on visual navigation methods based on deep reinforcement learning[J]. Computer Engineering and Applications, 2025, 61(10):66-78.(in Chinese)
[4] Choi D, Fung A, Wang H, et al. Find everything:a general vision language model approach to multi-object search[C] //2025 IEEE/RSJ International Conference on Intelligent Robots and Systems, 2025:19936-19943.
[5] Yokoyama N, Ha S, Batra D, et al. VLFM:vision-language frontier maps for zero-shot semantic navigation[C] //2024 IEEE International Conference on Robotics and Automation, 2024:42-48.
[6] Zhou K W, Zheng K Z, Pryor C, et al. ESC:exploration with soft commonsense constraints for zero-shot object navigation[C] //International Conference on Machine Learning, 2023:202.
[7] Wu P, Mu Y, Wu B, et al. VoroNav:voronoi-based zero-shot object navigation with large language model[PP/OL]. V2(2024-02-06)[2026-01-28]. https://doi.org/10.48550/arXiv.2401.02695.
[8] Gu J, Stefani E, Wu Q, et al. Vision-and-language navigation:a survey of tasks, methods, and future directions[C] //60th Annual Meeting of the Association for Computational Linguistics,2022:7606-7623.
[9] Shen W B, Xu D F, Zhu Y K, et al. Situational fusion of visual representation for visual navigation[C] //2019 IEEE/CVF International Conference on Computer Vision, 2019:2881-2890.
[10] Druon R, Yoshiyasu Y, Kanezaki A, et al. Visual object search by learning spatial context[J].IEEE Robotics and Automation Letters, 2020, 5(2):1279-1286.
[11] Ramakrishnan S K, Chaplot D S, Alhalah Z, et al. PONI:potential functions for objectgoal navigation with interaction-free learning[C] //2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022:18868-18878.
[12] Gervet T, Chintala S, Batra D, et al. Navigating to objects in the real world[J]. Science Robotics,2023, 8(79):eabm9798.
[13] Li L, Zuo X K, Peng H X, et al. Improving autonomous exploration using reduced approximated generalized voronoi graphs[J]. Journal of Intelligent&Robotic Systems, 2020, 99(1):91-113.
[14] Majumdar A, Aggarwal G, Devnani B, et al. ZSON:zero-shot object-goal navigation using multimodal goal embeddings[C] //36th International Conference on Neural Information Processing Systems, 2022:2343.
[15] Zhao Q F, Zhang L, He B, et al. Semantic policy network for zero-shot object goal visual navigation[J]. IEEE Robotics and Automation Letters, 2023. 8(11):7655-7662.
[16] 符腾藩,路飞,李忠阳,等.基于开放词汇目标检测的零样本目标导航方法[J/OL].机器人, 2025-08-27. https://doi.org/10.13973/j.cnki.robot.250180.Fu T F, Lu F, Li Z Y, et al. Zero-shot object navigation method based on open-vocabulary object detection[J/OL]. Robot, 2025-08-27. https://doi.org/10.13973/j.cnki.robot.250180.(in Chinese)
[17] Yu B, Kasaei H, Cao M. L3MVN:leveraging large language models for visual target navigation[C] //2023 IEEE/RSJ International Conference on Intelligent Robots and Systems, 2023:3554-3560.
[18] Armeni I, He Z Y, Gwak J, et al. 3D scene graph:a structure for unified semantics, 3D space,and camera[C] //IEEE International Conference on Computer Vision, 2019:5664-5673.
[19] Li J N, Li D X, Savarese S, et al. BLIP-2:bootstrapping language-image pre-training with frozen image encoders and large language models[C] //International Conference on Machine Learning,2023:202.
[20] Yamauchi B. A frontier-based approach for autonomous exploration[C] //IEEE International Symposium on Computational Intelligence in Robotics and Automation, 1997:146-151.
[21] Lin T Y, Maire M, Belongie S, et al. Microsoft COCO:common objects in context[C] //European Conference on Computer Vision, 2014, 8693:740-755.
[22] Wang C Y, Bochkovskiy A, Liao H Y M. YOLOv7:trainable bag-of-freebies sets new stateof-the-art for real-time object detectors[C] //2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023:7464-7475.
[23] Liu S, Zeng Z, Ren T, et al. Grounding dino:marrying dino with grounded pre-training for open-set object detection[C] //European Conference on Computer Vision, 2024, 15105:38-55.
[24] Zhang C, Han D, Qiao Y, et al. Faster segment anything:towards lightweight sam for mobile applications[PP/OL]. V2(2023-07-01)[2026-01-28]. https://doi.org/10.48550/arXiv.2306.14289.
[25] Wijmans E, Essa I, Batra D. VER:scaling on-policy RL leads to the emergence of navigation in embodied rearrangement[C] //Advances in Neural Information Processing Systems 35, 2022:7727-7740.
[26] Gadre S Y, Wortsman M, Ilharco G, et al. Cows on pasture:baselines and benchmarks for language-driven zero-shot object navigation[C] //2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023:23171-23181.
[27] Long Y X, Cai W Z, Wang H C, et al. Instructnav:zero-shot system for generic instruction navigation in unexplored environment[C] //Conference on Robot Learning, 2024:270.
Outlines

/