应用科学学报 ›› 2026, Vol. 44 ›› Issue (4): 626-643.doi: 10.3969/j.issn.0255-8297.2026.04.008

• 智能视觉感知 • 上一篇    下一篇

视觉感知的投影式三维语义分割后处理方法

邹国镜1, 丛铭1, 崔建军1, 韩玲2   

  1. 1. 长安大学 地质工程与测绘学院, 陕西 西安 710064;
    2. 长安大学 土地工程学院, 陕西 西安 710064
  • 收稿日期:2026-01-30 发布日期:2026-08-01
  • 通信作者: 丛铭,副教授,研究方向为自适应自学习、深度神经网络。E-mail:mingc@chd.edu.cn E-mail:mingc@chd.edu.cn
  • 基金资助:
    陕西省教育厅服务地方专项计划项目(No.23JE002)

Visual Perception-Based Post-Processing Method for Projection-Based Three-Dimensional Semantic Segmentation

ZOU Guojing1, CONG Ming1, CUI Jianjun1, HAN Ling2   

  1. 1. School of Geological Engineering and Geomatics, Chang'an University, Xi'an 710064, Shaanxi, China;
    2. School of Land Engineering, Chang'an University, Xi'an 710064, Shaanxi, China
  • Received:2026-01-30 Published:2026-08-01

摘要: 投影式三维语义分割方法能够有效降低三维数据处理复杂度与计算开销,但现有方法大多依赖RGB颜色空间进行特征提取,易受光照变化、阴影干扰以及类别颜色相似性的影响,导致分割精度受限。针对上述问题,本文受人类视觉系统颜色感知机制启发,提出一种基于多颜色空间后处理的投影式三维语义分割方法。首先,通过二维投影与栅格化处理将三维场景转换为二维规则影像,并利用深度学习模型获得初始分割结果;其次,在低置信度区域引入更符合人类视觉感知特性的Lab与HSV颜色空间特征进行聚类优化,以提升类别可分性与区域一致性;再次,结合二维–三维映射关系,将优化后的语义标签恢复至三维场景,实现高精度语义标注。实验结果表明,本文方法在4组复杂城市场景数据上均取得较好的分割效果。相较传统聚类方法,总体分类精度(overall accuracy,OA)平均提升约21%,Kappa系数平均提升约0.28;相较单纯深度学习方法,OA平均提升约5%,Kappa系数平均提升约0.07。结果表明,该方法能够有效提升复杂场景下三维语义分割的精度与稳定性。

关键词: 三维语义分割, 投影式方法, 后处理优化, 深度学习

Abstract: Projection-based three-dimensional semantic segmentation methods can effectively reduce the processing complexity and computational cost of three-dimensional data. However, most existing methods rely on the RGB color space for feature extraction, making them susceptible to illumination variations, shadow interference, and category color similarity, which limits segmentation accuracy. To address these issues, inspired by the color perception mechanism of the human visual system, this paper proposed a projection-based three-dimensional semantic segmentation method based on multi-color space post-processing. First, the three-dimensional scene was transformed into a regular two-dimensional image through two-dimensional projection and rasterization, and an initial segmentation result was obtained using a deep learning model. Second, in low-confidence regions, Lab and HSV color space features that are more consistent with human visual perception characteristics were introduced for clustering optimization to improve category separability and regional consistency. Third, combined with the two-dimensional–threedimensional mapping relationship, the optimized semantic labels were restored to the threedimensional scene, achieving high-precision semantic annotation. The experimental results show that the proposed method achieves good segmentation performance on four sets of complex urban scene data. Compared with traditional clustering methods, the overall accuracy(OA) is improved by approximately 21% on average, and the Kappa coefficient is increased by approximately 0.28 on average. Compared with deep learning-only methods,the OA is improved by approximately 5% on average, and the Kappa coefficient is increased by approximately 0.07 on average. The results indicate that the proposed method can effectively enhance the accuracy and stability of three-dimensional semantic segmentation in complex scenes.

Key words: three-dimensional semantic segmentation, projection-based method, postprocessing optimization, deep learning

中图分类号: