Journal of Applied Sciences ›› 2026, Vol. 44 ›› Issue (4): 615-625.doi: 10.3969/j.issn.0255-8297.2026.04.007

• Intelligent Visual Perception • Previous Articles     Next Articles

Zero-Shot Target Navigation Integrating Visual Scene Understanding and Commonsense Reasoning

ZHOU Feiyu1,2,3, ZHANG Xing1,2,3   

  1. 1. School of Architecture & Urban Planning, Shenzhen University, Shenzhen 518060, Guangdong, China;
    2. Guangdong Key Laboratory of Urban Informatics, Shenzhen University, Shenzhen 518060, Guangdong, China;
    3. Key Laboratory for Geo-Environmental Monitoring of Great Bay Area, MNR, Shenzhen University, Shenzhen 518060, Guangdong, China
  • Received:2026-01-30 Published:2026-08-01

Abstract: To address the problem that existing zero-shot navigation methods are difficult to simultaneously satisfy complex instruction understanding and efficient visual search, a zero-shot target navigation framework integrating the dual-level commonsense reasoning of large language models and the synergistic environmental perception of vision-language models was proposed. Through a modular design, the large language model was positioned as a commonsense reasoning engine to parse the target from natural language instructions and generate dual-level “room-object” spatial search priors. The vision-language model was taken as a real-time scene interpreter, which performed semantic matching based on this prior and output local semantic values. The two models collaborated to construct a semantic value map, achieving an organic unification of the commonsense-guided global search and the vision-driven local localization. Experimental results indicate that the proposed method improves the navigation success rate and the path efficiency. Its outstanding performance on navigation efficiency metrics validates the effectiveness of the spatial prior reasoning of large language models in improving search efficiency. Under the zero-shot setting, this framework achieves a deep integration of real-time perception and commonsense generalization, providing an efficient, accurate, and practical solution for intelligent navigation in open scenarios.

Key words: large language model, vision-language model, zero-shot target navigation, vision-language navigation, embodied intelligence

CLC Number: