应用科学学报 ›› 2026, Vol. 44 ›› Issue (4): 685-700.doi: 10.3969/j.issn.0255-8297.2026.04.012

• 智能视觉感知 • 上一篇    

基于空频感知与差分互补的语义驱动实时红外可见光融合

张组军1, 甘睿2, 李威1, 陈九九1, 谢宇1, 熊邦书1   

  1. 1. 南昌航空大学 信息工程学院, 江西 南昌 330063;
    2. 江西省交通投资集团有限责任公司, 江西 南昌 330013
  • 收稿日期:2026-02-05 发布日期:2026-08-01
  • 通信作者: 熊邦书,教授,研究方向为图像处理,计算机视觉和直升机故障诊断。E-mail:xiongbs@126.com E-mail:xiongbs@126.com
  • 基金资助:
    国家自然科学基金(No.62473187);江西省职业早期青年科技人才项目(20244BCE52091);江西省教育厅科技项目(GJJ2401015);江西省图像处理与模式识别重点实验室开放基金(ET202404438)资助

Semantic-Driven Real-Time Infrared and Visible Image Fusion Based on Spatial-Frequency Perception and Differential Complementation

ZHANG Zujun1, GAN Rui2, LI Wei1, CHEN Jiujiu1, XIE Yu1, XIONG Bangshu1   

  1. 1. School of Information Engineering, Nanchang Hangkong University, Nanchang 330063, Jiangxi, China;
    2. Jiangxi Communications Investment Group Co., Ltd., Nanchang 330013, Jiangxi, China
  • Received:2026-02-05 Published:2026-08-01

摘要: 针对现有方法全局特征提取受限、跨模态互补特征融合不充分,以及忽视下游视觉任务需求等问题,提出一种基于空频感知与差分互补的语义驱动实时红外可见光融合方法。首先,设计空频双分支感知模块,有效捕捉图像的局部空间细节与全局依赖关系;其次,设计跨模态差分特征互补模块,充分融合不同模态的优势特征;再次,构建语义驱动的联合训练框架,利用语义分割损失引导融合网络保留更多语义信息,从而提高下游高级视觉任务的性能。在RoadScene、MSRS及M3FD公开数据集上的实验结果表明,所提方法与主流方法相比,融合图像的互信息和视觉保真度分别平均提升了3.20%和3.79%;在语义分割任务中,平均交并比提升了2.98%;在目标检测任务中,平均精度均值提升了5.64%。在运行效率方面,处理帧率达到38.99 FPS,满足工程应用的实时性要求。

关键词: 图像融合, 空频双分支感知, 跨模态差分特征互补, 语义驱动

Abstract: To address the limitations of existing methods, such as limited global feature extraction, insufficient fusion of cross-modal complementary features, and the neglect of downstream visual task requirements, a semantic-driven real-time infrared and visible image fusion method based on spatial-frequency perception and differential complementation is proposed. First, a spatial-frequency dual-branch perception module was designed to effectively capture the local spatial details and global dependencies of the image. Second, a cross-modality differential feature complementation module was designed to fully integrate the advantageous features of different modalities. Furthermore, a semantic-driven joint training framework was constructed, using semantic segmentation loss to guide the fusion network to retain more semantic information, thereby improving the performance of downstream advanced visual tasks. Experimental results on the public datasets RoadScene, MSRS, and M3FD show that, compared with mainstream methods, the mutual information and visual fidelity of the fused images are improved by an average of 3.20% and 3.79%, respectively, by the proposed method; in the semantic segmentation task, the mean intersection over union is improved by 2.98%; in the object detection task, the mean average precision is improved by 5.64%. In terms of running efficiency, the processing frame rate reaches 38.99 FPS, meeting the real-time requirements of engineering applications.

Key words: image fusion, spatial-frequency dual-branch perception, cross-modality differential feature complementation, semantic drive

中图分类号: