• Home
  • Journal
  • Editorial Board
  • Instruction
  • Subscription
  • Solicit
  • Contact Us
中文
Picture News
Previous Next
Featured Articles More>>
News More>>
  • 1 About G2020-0098 (2020-06-13)
  • 2 New Submition Notes (2010-07-07)
  • 3 《应用科学学报》第二届编委会会议在上海大学召开 (2009-03-11)
Current IssueMore>>
31 July 2026, Volume 44 Issue 4
Previous Issue   
Intelligent Visual Perception
Review of LiDAR Simultaneous Localization and Mapping
DUAN Xuzhe, ZHONG Ruofei, LI Jian, FU Jing, ZHAO Pengcheng, LI Jiayuan, AI Mingyao, HU Qingwu
2026, 44(4):  515-536.  doi:10.3969/j.issn.0255-8297.2026.04.001
Asbtract ( 19 )   PDF (1771KB) ( 5 )  
References | Related Articles | Metrics
The development of light detection and ranging(LiDAR) simultaneous localization and mapping(SLAM) technology over the past decade was systematically reviewed.First, a general mathematical definition of the SLAM problem was provided to establish a unified analytical framework. Subsequently, the related research was categorized into two main categories according to the number of platforms: single-platform LiDAR SLAM and multi-platform collaborative LiDAR SLAM. In the single-platform section, the multi-sensor fusion methods centered on LiDAR were emphatically summarized; in the multi-platform section, the currently mature collaborative SLAM systems and their key characteristics were outlined. Meanwhile, the challenges and development opportunities faced by LiDAR SLAM in complex scenes, degraded environments, multi-modal data sources, and multiplatform collaboration were discussed. Finally, combined with current research hotspots,the future technical evolution trends of LiDAR SLAM were prospected. The main application forms of deep learning in LiDAR SLAM were analyzed, and the potential value of emerging map representations(such as neural radiance fields and three-dimensional Gaussian splatting) in mapping quality and representation capability was outlined.
Construction of Multi-sensor Simulation System for Laser SLAM
LI Kaixin, XU Zhihong, SUN Zhenxing, ZHONG Ruofei
2026, 44(4):  537-552.  doi:10.3969/j.issn.0255-8297.2026.04.002
Asbtract ( 22 )   PDF (3699KB) ( 5 )  
References | Related Articles | Metrics
Laser SLAM digital simulation provides an efficient, controllable, and reproducible virtual test environment for the development of mobile measurement systems, significantly reducing the reliance on complex field tests. However, existing methods mostly rely on three-dimensional modeling, which face problems such as high modeling costs, geometric distortion, and difficulty in synchronously generating multi-sensor data. Therefore,a method for constructing a digital simulation system for mobile measurement based on laser simultaneous localization and mapping(SLAM) directly driven by real point clouds was proposed in this paper. The original laser point cloud data were directly transformed into a virtual environment with real scene characteristics. The behaviors and state changes of the sensors on the acquisition trajectory were simulated through the inverse digital modeling of the core sensors of laser radar and inertial measurement unit and the adoption of sparse voxel grid index and ray casting algorithm. The system supported the independent planning of B-spline curve trajectories and the flexible configuration of parameters such as scanning resolution and field of view and could synchronously generate spatiotemporally consistent raw laser radar data and inertial measurement unit data. Experimental results show that compared with traditional modeling-based simulation methods, the proposed method significantly reduces the cost and technical threshold of scene construction. Meanwhile, under the same scene and acquisition trajectory conditions, the geometric errors of point clouds, namely root mean square error(RMSE) and mean absolute error(MAE),decrease by approximately 18%~33% and 40%~48%, respectively, compared with the comparison methods. This indicates that the proposed method can effectively solve the high threshold problem of traditional modeling and provides a virtual test means with high geometric fidelity for SLAM algorithm evaluation and equipment development.
Visibility Evaluation of Scenic Spots in Mountain Heritage Parks Based on LiDAR Data: A Case Study of Baoxie Plank Road Heritage Park
LI Yifan, Anna Mária Csergő, CHEN Zhipeng, LIANG Anbang, LI Yongrong, ZHANG Xujie
2026, 44(4):  553-570.  doi:10.3969/j.issn.0255-8297.2026.04.003
Asbtract ( 14 )   PDF (10144KB) ( 35 )  
References | Related Articles | Metrics
Visibility design of heritage parks is the key to implementing the planning concept of “experience-oriented”, which directly affects the dissemination of cultural heritage values. In view of the difficulty of traditional methods in accurately capturing the threedimensional spatial characteristics of mountain sites, a visibility analysis and evaluation system based on LiDAR data was established by taking the Baoxie Plank Road Heritage Park as an example. A technical framework of “data acquisition, preprocessing, model construction, viewshed analysis, quantitative evaluation, and optimization suggestions” was adopted. Data were collected by a backpack LiDAR system, and a three-dimensional voxel model with a precision of ≤ 5 cm was constructed. On this basis, three core indicators,including three-dimensional visibility, visual openness, and the proportion of ground object types in the view, were selected to carry out a quantitative analysis. The results show that there are significant differences in visibility among the three core scenic spots; the plank beam scenic spot H1 has the best visibility(the three-dimensional visibility of the optimal viewpoint is 1.000); the plank hole scenic spot Z1 is heavily blocked by vegetation and terrain, and its average three-dimensional visibility is only 0.044; the visibility of the scenic spot Z2 shows a significant spatial differentiation, with the average value of its near viewpoints reaching 0.708, that of the far viewpoints being only 0.056, and the overall average value being 0.251. In addition, terrain, vegetation(including tall arbors and low shrubs), heritage location, and viewing facilities are the four key factors affecting visibility.The optimization suggestions proposed in this paper can provide technical support for the planning of heritage parks and similar studies.
Speech Translation Method Enhanced with Context Awareness and Modality Alignment
YANG Songtao, GAO Shengxiang, OUYANG Fei, YU Zhengtao
2026, 44(4):  571-584.  doi:10.3969/j.issn.0255-8297.2026.04.004
Asbtract ( 41 )   PDF (815KB) ( 19 )  
References | Related Articles | Metrics
To address the problem of insufficient perception of contextual information in speech representations by modality alignment modules in speech translation tasks, a contextual semantic-enhanced speech-to-text translation method was proposed. Building upon speech representations, a contextual semantic reorganization mechanism based on local window variance perception and similarity measurement was introduced to perform semantic-aware representation compression on frame-level speech representations, so as to obtain final speech representations with contextual semantic enhancement. Meanwhile, a global knowledge modality adapter was introduced to further achieve modality alignment between the speech contextual representation and the text. Finally, the fused multimodal representation was input to a large speech model decoder, which decoded it to generate the translated text of the target language. Experimental results demonstrate that the method effectively extracts speech representations containing contextual semantic information and achieves better performance in downstream translation tasks.
A Disease Detection Method for Point Cloud of Stone Cultural Relics Based on GMM
ZHANG Jiamin, CUI Hao, LYU Hongyi, JIAO Jianhui
2026, 44(4):  585-597.  doi:10.3969/j.issn.0255-8297.2026.04.005
Asbtract ( 22 )   PDF (2453KB) ( 6 )  
References | Related Articles | Metrics
Stone cultural relics are important material carriers of civilization, and detecting their surface damage is a core task in the preventive conservation of cultural heritage.However, when dealing with point cloud data characterized by complex surface textures and varying degrees of weathering, existing detection and segmentation algorithms exhibit two limitations: 1) traditional hard-threshold methods rely on manual parameter tuning and exhibit poor adaptability; 2) artificial textures and natural weathering damage are highly similar in geometric features, which leads to region merging and confusion between them during segmentation. Therefore, an automatic detection method combining a statistical model and geometric topological analysis was proposed. First, geometric roughness and color consistency were selected to construct a two-dimensional feature vector, and unsupervised adaptive anomaly extraction was realized by a Gaussian mixture model to replace the traditional hard-threshold method relying on manual parameter tuning. Subsequently, a multi-constraint locally convex connected patches(LCCP) segmentation algorithm was designed. By introducing an orthogonal distance constraint and a normal angle constraint, it solved the region-merging problem between textures and damage from a geometric-topological perspective, and a post-processing step was conducted to optimize the segmentation results. Validation was carried out on three typical test regions with texture–damage merging. Compared with the traditional LCCP algorithm and region growing algorithm, indicators of the proposed method such as precision, F1-score, and intersection over union were comprehensively improved; the F1-scores on the three test regions reached 0.915 8, 0.906 3, and 0.926 0, respectively. The proposed method effectively resolved the merging problem between textures and damage, and provided a reliable solution for the digital detection of damage in stone cultural relics.
Methods and Applications of Small Object Detection in Remote Sensing Images
WANG Yufan, SHAO Zilong, PU Yuanxue, ZHANG Haowen, ZHANG Jingyi, QIN Kun
2026, 44(4):  598-614.  doi:10.3969/j.issn.0255-8297.2026.04.006
Asbtract ( 22 )   PDF (5886KB) ( 12 )  
References | Related Articles | Metrics
Small object detection in remote sensing imagery is one of the core issues in the fields of remote sensing and computer vision, playing a crucial role in the intelligent understanding of remote sensing scenes and intelligent interpretation of remote sensing images. Due to factors such as the low pixel ratio of targets and complex scenes, small object detection in remote sensing imagery faces multiple challenges, including difficulties in target feature extraction, susceptibility to noise and occlusion interference, stringent requirements for bounding box regression accuracy, scarcity of dedicated datasets, and insufficient model design adaptability. The technical difficulties and bottlenecks of small object detection in remote sensing imagery were systematically reviewed. Six mainstream detection methods, including traditional remote sensing image processing detection methods, dataset augmentation-based detection methods, resolution enhancement-based detection methods,feature fusion and attention mechanism-based detection methods, Transformer-based detection methods, and anchor-free technology-based detection methods, were analyzed, and the technical principles, core innovations, and applicable scenarios of each method were detailed. The characteristics, annotation specifications, and application value of remote sensing small object datasets were summarized, and the applications of remote sensing small object detection technology in ship and navigation mark detection, vehicle detection,and military target detection were analyzed. Finally, future development directions for remote sensing small object detection, including multimodal fusion, lightweight architecture,and few-shot learning, were prospected.
Zero-Shot Target Navigation Integrating Visual Scene Understanding and Commonsense Reasoning
ZHOU Feiyu, ZHANG Xing
2026, 44(4):  615-625.  doi:10.3969/j.issn.0255-8297.2026.04.007
Asbtract ( 10 )   PDF (4162KB) ( 7 )  
References | Related Articles | Metrics
To address the problem that existing zero-shot navigation methods are difficult to simultaneously satisfy complex instruction understanding and efficient visual search, a zero-shot target navigation framework integrating the dual-level commonsense reasoning of large language models and the synergistic environmental perception of vision-language models was proposed. Through a modular design, the large language model was positioned as a commonsense reasoning engine to parse the target from natural language instructions and generate dual-level “room-object” spatial search priors. The vision-language model was taken as a real-time scene interpreter, which performed semantic matching based on this prior and output local semantic values. The two models collaborated to construct a semantic value map, achieving an organic unification of the commonsense-guided global search and the vision-driven local localization. Experimental results indicate that the proposed method improves the navigation success rate and the path efficiency. Its outstanding performance on navigation efficiency metrics validates the effectiveness of the spatial prior reasoning of large language models in improving search efficiency. Under the zero-shot setting, this framework achieves a deep integration of real-time perception and commonsense generalization, providing an efficient, accurate, and practical solution for intelligent navigation in open scenarios.
Visual Perception-Based Post-Processing Method for Projection-Based Three-Dimensional Semantic Segmentation
ZOU Guojing, CONG Ming, CUI Jianjun, HAN Ling
2026, 44(4):  626-643.  doi:10.3969/j.issn.0255-8297.2026.04.008
Asbtract ( 12 )   PDF (9141KB) ( 5 )  
References | Related Articles | Metrics
Projection-based three-dimensional semantic segmentation methods can effectively reduce the processing complexity and computational cost of three-dimensional data. However, most existing methods rely on the RGB color space for feature extraction, making them susceptible to illumination variations, shadow interference, and category color similarity, which limits segmentation accuracy. To address these issues, inspired by the color perception mechanism of the human visual system, this paper proposed a projection-based three-dimensional semantic segmentation method based on multi-color space post-processing. First, the three-dimensional scene was transformed into a regular two-dimensional image through two-dimensional projection and rasterization, and an initial segmentation result was obtained using a deep learning model. Second, in low-confidence regions, Lab and HSV color space features that are more consistent with human visual perception characteristics were introduced for clustering optimization to improve category separability and regional consistency. Third, combined with the two-dimensional–threedimensional mapping relationship, the optimized semantic labels were restored to the threedimensional scene, achieving high-precision semantic annotation. The experimental results show that the proposed method achieves good segmentation performance on four sets of complex urban scene data. Compared with traditional clustering methods, the overall accuracy(OA) is improved by approximately 21% on average, and the Kappa coefficient is increased by approximately 0.28 on average. Compared with deep learning-only methods,the OA is improved by approximately 5% on average, and the Kappa coefficient is increased by approximately 0.07 on average. The results indicate that the proposed method can effectively enhance the accuracy and stability of three-dimensional semantic segmentation in complex scenes.
Three-Dimensional Scene Graph Representation of Forest Point Clouds and LLM Applications
PENG Jinhong, CHEN Maolin, XUE Mei, LI Rufeng
2026, 44(4):  644-656.  doi:10.3969/j.issn.0255-8297.2026.04.009
Asbtract ( 8 )   PDF (2843KB) ( 4 )  
References | Related Articles | Metrics
Three-dimensional laser scanning technology has become a crucial method for obtaining high-precision forest stand parameters in forestry surveys. However, the expression of forest three-dimensional scenes based on massive laser point clouds relies on professional software for processing, and it primarily focuses on object-level semantic understanding, lacking explicit expression of element relationship information. To address this issue, a three-dimensional scene graph conceptual model oriented to forest scenes was proposed, and the methods of element hierarchical classification, geometric description,semantic expression, and relationship description of forest scenes were systematically introduced. On this basis, the three-dimensional scene graph was combined with semantic association analysis, and a large language model was introduced as a query and analysis tool. A three-level universal evaluation framework was designed to comparatively evaluate ChatGPT-4o, DeepSeek-R1, and Grok4 in terms of operating speed, input limitations, and visual effects, and the efficacy, potential, and limitations of this model were discussed. Experimental results based on the public point cloud dataset ForestSemantic show that the three-dimensional scene graph has a good data-carrying capacity in forest scenes and can effectively organize and connect forest relationships; the accuracy of Grok4 is above 90%,which is superior to the other two large language models. This study can provide assistance for the spatial expression and analysis, operation management, and decision-making of forestry resources.
Full-Reference Screen Content Image Quality Assessment via Text-Visual Fusion
ZHANG Yin, YANG Chao, AN Ping, HUANG Xinpeng
2026, 44(4):  657-668.  doi:10.3969/j.issn.0255-8297.2026.04.010
Asbtract ( 13 )   PDF (714KB) ( 3 )  
References | Related Articles | Metrics
With the rapid development of Internet technology, the application of screen content image(SCI) in the network is becoming increasingly widespread, and the issue of objective quality assessment has become a research hotspot. A full-reference SCI quality assessment method integrating text integrity and image perception features was proposed. To address the rich text characteristics of SCI, optical character recognition(OCR) technology was utilized to extract text content, and character error rate(CER) was calculated as the text integrity feature. Simultaneously, multiple quality perception features, including spatial, structural, and information-theoretic features, were incorporated. Random forest was employed to learn the adaptive weights of the features, and support vector regression(SVR) was utilized to establish the mapping relationship between the fused features and subjective quality scores. The experimental results on two widely used datasets, SCID and SIQAD, demonstrate that the proposed method achieves superior performance. Compared with existing competitive methods, the prediction accuracy of the proposed method on the SCID dataset improves by more than 2%, which more accurately reflects the perception characteristics of the human visual system for screen content.
Mamba-Based Pattern-Aware Gait Detection Algorithm and Application in Pedestrian Dead Reckoning
ZHOU Ziyi, LIU Deer, ZHONG Chonglin, WANG Yuan, ZHANG Tao, QIU Yongkang
2026, 44(4):  669-684.  doi:10.3969/j.issn.0255-8297.2026.04.011
Asbtract ( 19 )   PDF (3890KB) ( 3 )  
References | Related Articles | Metrics
To address the issues of missed and false detections in gait detection algorithms under complex motion modes and varying device carrying postures, an adaptive gait detection algorithm integrating a Mamba neural network and a finite state machine(FSM) was proposed. A classification model based on the Mamba network was constructed to jointly identify pedestrian motion modes and smartphone carrying postures. Furthermore, this algorithm fused an adaptive-threshold FSM, an angular velocity backtracking mechanism,and a gravity projection-based non-rigid coupling denoising strategy to achieve the stable triggering of gait events under complex carrying conditions, thereby further improving the positioning accuracy and robustness of pedestrian dead reckoning(PDR) in complex scenarios. Experimental results demonstrate that the proposed algorithm achieves high gait detection accuracy under various motion modes and smartphone carrying postures,effectively suppresses false and missed detections, and reduces the PDR positioning errors introduced by the gait detection stage.
Semantic-Driven Real-Time Infrared and Visible Image Fusion Based on Spatial-Frequency Perception and Differential Complementation
ZHANG Zujun, GAN Rui, LI Wei, CHEN Jiujiu, XIE Yu, XIONG Bangshu
2026, 44(4):  685-700.  doi:10.3969/j.issn.0255-8297.2026.04.012
Asbtract ( 20 )   PDF (27713KB) ( 11 )  
References | Related Articles | Metrics
To address the limitations of existing methods, such as limited global feature extraction, insufficient fusion of cross-modal complementary features, and the neglect of downstream visual task requirements, a semantic-driven real-time infrared and visible image fusion method based on spatial-frequency perception and differential complementation is proposed. First, a spatial-frequency dual-branch perception module was designed to effectively capture the local spatial details and global dependencies of the image. Second, a cross-modality differential feature complementation module was designed to fully integrate the advantageous features of different modalities. Furthermore, a semantic-driven joint training framework was constructed, using semantic segmentation loss to guide the fusion network to retain more semantic information, thereby improving the performance of downstream advanced visual tasks. Experimental results on the public datasets RoadScene, MSRS, and M3FD show that, compared with mainstream methods, the mutual information and visual fidelity of the fused images are improved by an average of 3.20% and 3.79%, respectively, by the proposed method; in the semantic segmentation task, the mean intersection over union is improved by 2.98%; in the object detection task, the mean average precision is improved by 5.64%. In terms of running efficiency, the processing frame rate reaches 38.99 FPS, meeting the real-time requirements of engineering applications.
Office Online
Authors Login
Peer Review
Editorial Work
Editor-in-Chief
Office Work
Journal
  • Just Accepted
  • Current Issue
  • Archive
  • Advanced Search
  • Volumn Content
  • Most Read
  • Most Download
  • E-mail Alert
  • RSS
Download >
Links>
  • JAS E-mail
  • CNKI-check
  • SHU
Information
Bimonthly, Founded in 1983
Editor-in-Chief:Wang Tingyun
ISSN 0255-8297
CN 31-1404/N

Copyright © Editorial office of Journal of Applied Sciences

Tel: 86-21-66131736 
E-mail: yykxxb@department.shu.edu.cn