应用科学学报 ›› 2026, Vol. 44 ›› Issue (4): 615-625.doi: 10.3969/j.issn.0255-8297.2026.04.007

• 智能视觉感知 • 上一篇    下一篇

融合视觉场景理解与常识推理的零样本目标导航

周飞宇1,2,3, 张星1,2,3   

  1. 1. 深圳大学 建筑与城市规划学院, 广东 深圳 518060;
    2. 深圳大学 广东省城市空间信息工程重点实验室, 广东 深圳 518060;
    3. 深圳大学 自然资源部大湾区地理环境监测重点实验室, 广东 深圳 518060
  • 收稿日期:2026-01-30 发布日期:2026-08-01
  • 通信作者: 张星,博士,副教授,研究方向为无人系统自主导航与三维建图。E-mail:xzhang@szu.edu.cn E-mail:xzhang@szu.edu.cn
  • 基金资助:
    国家自然科学基金(No.42571483,No.42071434);广东省基础与应用基础研究基金(No.2025A1515011725);深圳市科技计划资助(No.JCYJ20240813142833044)

Zero-Shot Target Navigation Integrating Visual Scene Understanding and Commonsense Reasoning

ZHOU Feiyu1,2,3, ZHANG Xing1,2,3   

  1. 1. School of Architecture & Urban Planning, Shenzhen University, Shenzhen 518060, Guangdong, China;
    2. Guangdong Key Laboratory of Urban Informatics, Shenzhen University, Shenzhen 518060, Guangdong, China;
    3. Key Laboratory for Geo-Environmental Monitoring of Great Bay Area, MNR, Shenzhen University, Shenzhen 518060, Guangdong, China
  • Received:2026-01-30 Published:2026-08-01

摘要: 针对现有零样本导航方法难以同时满足复杂指令理解与高效视觉搜索的问题,提出一种融合大语言模型双层级常识推理与视觉语言模型环境感知协同的零样本目标导航框架。通过模块化设计,将大语言模型定位为常识推理机,从自然语言指令中解析目标并生成“房间-物体”双层级空间搜索先验。将视觉语言模型作为实时场景解释器,依据该先验进行语义匹配并输出局部语义价值。二者协同构建语义价值地图,实现常识引导的全局搜索与视觉驱动的局部定位的有机统一。实验结果表明,所提方法在导航成功率与路径效率上有所提升。在导航效率指标上表现突出,验证了大语言模型空间先验推理对提升搜索效率的有效性。该框架在零样本设定下实现了感知的实时性与常识泛化性的深度结合,为开放场景下的智能导航提供了兼具高效性、准确性与实用性的解决方案。

关键词: 大语言模型, 视觉语言模型, 零样本目标导航, 视觉语言导航, 具身智能

Abstract: To address the problem that existing zero-shot navigation methods are difficult to simultaneously satisfy complex instruction understanding and efficient visual search, a zero-shot target navigation framework integrating the dual-level commonsense reasoning of large language models and the synergistic environmental perception of vision-language models was proposed. Through a modular design, the large language model was positioned as a commonsense reasoning engine to parse the target from natural language instructions and generate dual-level “room-object” spatial search priors. The vision-language model was taken as a real-time scene interpreter, which performed semantic matching based on this prior and output local semantic values. The two models collaborated to construct a semantic value map, achieving an organic unification of the commonsense-guided global search and the vision-driven local localization. Experimental results indicate that the proposed method improves the navigation success rate and the path efficiency. Its outstanding performance on navigation efficiency metrics validates the effectiveness of the spatial prior reasoning of large language models in improving search efficiency. Under the zero-shot setting, this framework achieves a deep integration of real-time perception and commonsense generalization, providing an efficient, accurate, and practical solution for intelligent navigation in open scenarios.

Key words: large language model, vision-language model, zero-shot target navigation, vision-language navigation, embodied intelligence

中图分类号: