智能视觉感知

融合视觉场景理解与常识推理的零样本目标导航

  • 周飞宇 ,
  • 张星
展开
  • 1. 深圳大学 建筑与城市规划学院, 广东 深圳 518060;
    2. 深圳大学 广东省城市空间信息工程重点实验室, 广东 深圳 518060;
    3. 深圳大学 自然资源部大湾区地理环境监测重点实验室, 广东 深圳 518060

收稿日期: 2026-01-30

  网络出版日期: 2026-08-01

基金资助

国家自然科学基金(No.42571483,No.42071434);广东省基础与应用基础研究基金(No.2025A1515011725);深圳市科技计划资助(No.JCYJ20240813142833044)

Zero-Shot Target Navigation Integrating Visual Scene Understanding and Commonsense Reasoning

  • ZHOU Feiyu ,
  • ZHANG Xing
Expand
  • 1. School of Architecture & Urban Planning, Shenzhen University, Shenzhen 518060, Guangdong, China;
    2. Guangdong Key Laboratory of Urban Informatics, Shenzhen University, Shenzhen 518060, Guangdong, China;
    3. Key Laboratory for Geo-Environmental Monitoring of Great Bay Area, MNR, Shenzhen University, Shenzhen 518060, Guangdong, China

Received date: 2026-01-30

  Online published: 2026-08-01

摘要

针对现有零样本导航方法难以同时满足复杂指令理解与高效视觉搜索的问题,提出一种融合大语言模型双层级常识推理与视觉语言模型环境感知协同的零样本目标导航框架。通过模块化设计,将大语言模型定位为常识推理机,从自然语言指令中解析目标并生成“房间-物体”双层级空间搜索先验。将视觉语言模型作为实时场景解释器,依据该先验进行语义匹配并输出局部语义价值。二者协同构建语义价值地图,实现常识引导的全局搜索与视觉驱动的局部定位的有机统一。实验结果表明,所提方法在导航成功率与路径效率上有所提升。在导航效率指标上表现突出,验证了大语言模型空间先验推理对提升搜索效率的有效性。该框架在零样本设定下实现了感知的实时性与常识泛化性的深度结合,为开放场景下的智能导航提供了兼具高效性、准确性与实用性的解决方案。

本文引用格式

周飞宇 , 张星 . 融合视觉场景理解与常识推理的零样本目标导航[J]. 应用科学学报, 2026 , 44(4) : 615 -625 . DOI: 10.3969/j.issn.0255-8297.2026.04.007

Abstract

To address the problem that existing zero-shot navigation methods are difficult to simultaneously satisfy complex instruction understanding and efficient visual search, a zero-shot target navigation framework integrating the dual-level commonsense reasoning of large language models and the synergistic environmental perception of vision-language models was proposed. Through a modular design, the large language model was positioned as a commonsense reasoning engine to parse the target from natural language instructions and generate dual-level “room-object” spatial search priors. The vision-language model was taken as a real-time scene interpreter, which performed semantic matching based on this prior and output local semantic values. The two models collaborated to construct a semantic value map, achieving an organic unification of the commonsense-guided global search and the vision-driven local localization. Experimental results indicate that the proposed method improves the navigation success rate and the path efficiency. Its outstanding performance on navigation efficiency metrics validates the effectiveness of the spatial prior reasoning of large language models in improving search efficiency. Under the zero-shot setting, this framework achieves a deep integration of real-time perception and commonsense generalization, providing an efficient, accurate, and practical solution for intelligent navigation in open scenarios.

参考文献

[1] Batra D, Gokaslan A, Kembhavi A, et al. Objectnav revisited:on evaluation of embodied agents navigating to objects[PP/OL]. V2(2020-08-30)[2026-01-30].https://doi.org/10.48550/arXiv.2006.13171.
[2] 谢远龙,王书亭,程祥,等.大模型驱动的具身智能机器人导航技术综述[J].华中师范大学学报(自然科学版), 2025, 59(5):677-693.Xie Y L, Wang S T, Cheng X, et al. An overview of large model-driven embodied intelligent navigation[J]. Journal of Central China Normal University(Natural Sciences), 2025, 59(5):677-693.(in Chinese)
[3] 高宇宁,王安成,赵华凯,等.基于深度强化学习的视觉导航方法综述[J].计算机工程与应用, 2025,61(10):66-78.Gao Y N, Wang A C, Zhao H K, et al. Review on visual navigation methods based on deep reinforcement learning[J]. Computer Engineering and Applications, 2025, 61(10):66-78.(in Chinese)
[4] Choi D, Fung A, Wang H, et al. Find everything:a general vision language model approach to multi-object search[C] //2025 IEEE/RSJ International Conference on Intelligent Robots and Systems, 2025:19936-19943.
[5] Yokoyama N, Ha S, Batra D, et al. VLFM:vision-language frontier maps for zero-shot semantic navigation[C] //2024 IEEE International Conference on Robotics and Automation, 2024:42-48.
[6] Zhou K W, Zheng K Z, Pryor C, et al. ESC:exploration with soft commonsense constraints for zero-shot object navigation[C] //International Conference on Machine Learning, 2023:202.
[7] Wu P, Mu Y, Wu B, et al. VoroNav:voronoi-based zero-shot object navigation with large language model[PP/OL]. V2(2024-02-06)[2026-01-28]. https://doi.org/10.48550/arXiv.2401.02695.
[8] Gu J, Stefani E, Wu Q, et al. Vision-and-language navigation:a survey of tasks, methods, and future directions[C] //60th Annual Meeting of the Association for Computational Linguistics,2022:7606-7623.
[9] Shen W B, Xu D F, Zhu Y K, et al. Situational fusion of visual representation for visual navigation[C] //2019 IEEE/CVF International Conference on Computer Vision, 2019:2881-2890.
[10] Druon R, Yoshiyasu Y, Kanezaki A, et al. Visual object search by learning spatial context[J].IEEE Robotics and Automation Letters, 2020, 5(2):1279-1286.
[11] Ramakrishnan S K, Chaplot D S, Alhalah Z, et al. PONI:potential functions for objectgoal navigation with interaction-free learning[C] //2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022:18868-18878.
[12] Gervet T, Chintala S, Batra D, et al. Navigating to objects in the real world[J]. Science Robotics,2023, 8(79):eabm9798.
[13] Li L, Zuo X K, Peng H X, et al. Improving autonomous exploration using reduced approximated generalized voronoi graphs[J]. Journal of Intelligent&Robotic Systems, 2020, 99(1):91-113.
[14] Majumdar A, Aggarwal G, Devnani B, et al. ZSON:zero-shot object-goal navigation using multimodal goal embeddings[C] //36th International Conference on Neural Information Processing Systems, 2022:2343.
[15] Zhao Q F, Zhang L, He B, et al. Semantic policy network for zero-shot object goal visual navigation[J]. IEEE Robotics and Automation Letters, 2023. 8(11):7655-7662.
[16] 符腾藩,路飞,李忠阳,等.基于开放词汇目标检测的零样本目标导航方法[J/OL].机器人, 2025-08-27. https://doi.org/10.13973/j.cnki.robot.250180.Fu T F, Lu F, Li Z Y, et al. Zero-shot object navigation method based on open-vocabulary object detection[J/OL]. Robot, 2025-08-27. https://doi.org/10.13973/j.cnki.robot.250180.(in Chinese)
[17] Yu B, Kasaei H, Cao M. L3MVN:leveraging large language models for visual target navigation[C] //2023 IEEE/RSJ International Conference on Intelligent Robots and Systems, 2023:3554-3560.
[18] Armeni I, He Z Y, Gwak J, et al. 3D scene graph:a structure for unified semantics, 3D space,and camera[C] //IEEE International Conference on Computer Vision, 2019:5664-5673.
[19] Li J N, Li D X, Savarese S, et al. BLIP-2:bootstrapping language-image pre-training with frozen image encoders and large language models[C] //International Conference on Machine Learning,2023:202.
[20] Yamauchi B. A frontier-based approach for autonomous exploration[C] //IEEE International Symposium on Computational Intelligence in Robotics and Automation, 1997:146-151.
[21] Lin T Y, Maire M, Belongie S, et al. Microsoft COCO:common objects in context[C] //European Conference on Computer Vision, 2014, 8693:740-755.
[22] Wang C Y, Bochkovskiy A, Liao H Y M. YOLOv7:trainable bag-of-freebies sets new stateof-the-art for real-time object detectors[C] //2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023:7464-7475.
[23] Liu S, Zeng Z, Ren T, et al. Grounding dino:marrying dino with grounded pre-training for open-set object detection[C] //European Conference on Computer Vision, 2024, 15105:38-55.
[24] Zhang C, Han D, Qiao Y, et al. Faster segment anything:towards lightweight sam for mobile applications[PP/OL]. V2(2023-07-01)[2026-01-28]. https://doi.org/10.48550/arXiv.2306.14289.
[25] Wijmans E, Essa I, Batra D. VER:scaling on-policy RL leads to the emergence of navigation in embodied rearrangement[C] //Advances in Neural Information Processing Systems 35, 2022:7727-7740.
[26] Gadre S Y, Wortsman M, Ilharco G, et al. Cows on pasture:baselines and benchmarks for language-driven zero-shot object navigation[C] //2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023:23171-23181.
[27] Long Y X, Cai W Z, Wang H C, et al. Instructnav:zero-shot system for generic instruction navigation in unexplored environment[C] //Conference on Robot Learning, 2024:270.
文章导航

/