针对现有方法全局特征提取受限、跨模态互补特征融合不充分,以及忽视下游视觉任务需求等问题,提出一种基于空频感知与差分互补的语义驱动实时红外可见光融合方法。首先,设计空频双分支感知模块,有效捕捉图像的局部空间细节与全局依赖关系;其次,设计跨模态差分特征互补模块,充分融合不同模态的优势特征;再次,构建语义驱动的联合训练框架,利用语义分割损失引导融合网络保留更多语义信息,从而提高下游高级视觉任务的性能。在RoadScene、MSRS及M3FD公开数据集上的实验结果表明,所提方法与主流方法相比,融合图像的互信息和视觉保真度分别平均提升了3.20%和3.79%;在语义分割任务中,平均交并比提升了2.98%;在目标检测任务中,平均精度均值提升了5.64%。在运行效率方面,处理帧率达到38.99 FPS,满足工程应用的实时性要求。
To address the limitations of existing methods, such as limited global feature extraction, insufficient fusion of cross-modal complementary features, and the neglect of downstream visual task requirements, a semantic-driven real-time infrared and visible image fusion method based on spatial-frequency perception and differential complementation is proposed. First, a spatial-frequency dual-branch perception module was designed to effectively capture the local spatial details and global dependencies of the image. Second, a cross-modality differential feature complementation module was designed to fully integrate the advantageous features of different modalities. Furthermore, a semantic-driven joint training framework was constructed, using semantic segmentation loss to guide the fusion network to retain more semantic information, thereby improving the performance of downstream advanced visual tasks. Experimental results on the public datasets RoadScene, MSRS, and M3FD show that, compared with mainstream methods, the mutual information and visual fidelity of the fused images are improved by an average of 3.20% and 3.79%, respectively, by the proposed method; in the semantic segmentation task, the mean intersection over union is improved by 2.98%; in the object detection task, the mean average precision is improved by 5.64%. In terms of running efficiency, the processing frame rate reaches 38.99 FPS, meeting the real-time requirements of engineering applications.
[1] 金安安, 李祥, 张丽, 等. 基于NSCT与压缩感知的红外影像融合[J]. 应用科学学报, 2022, 40(1): 80-92. Jin A A, Li X, Zhang L, et al. Infrared image fusion based on NSCT and compressed sensing [J]. Journal of Applied Sciences, 2022, 40(1): 80-92. (in Chinese)
[2] Li Y, Li D, Fang S, et al. Multimodal image fusion network with prior-guided dynamic degradation removal for extreme environment perception [J]. Scientific Reports, 2025, 15(1): 40530- 40551.
[3] Li S, Yang B, Hu J. Performance comparison of different multi-resolution transforms for image fusion [J]. Information Fusion, 2011, 12(2): 74-84.
[4] Li H, Wu X J. Multi-focus image fusion using dictionary learning and low-rank representation [C] //International Conference on Image and Graphics, 2017: 675-686.
[5] Li S, Yin H, Fang L. Group-sparse representation with dictionary learning for medical image denoising and fusion [J]. IEEE Transactions on Biomedical Engineering, 2012, 59(12): 3450-3459.
[6] Ma J, Tang L, Xu M, et al. STDFusionNet: an infrared and visible image fusion network based on salient target detection [J]. IEEE Transactions on Instrumentation and Measurement, 2021, 70: 1-13.
[7] Xu J, Zhang J. Infrared and visible image fusion based on Haar wavelet downsampling and multi-scale feature aggregation [J]. Signal, Image and Video Processing, 2025, 19(14): 1200- 1208.
[8] Duan X, Liu Y, Yang Z, et al. CCFusion: a novel channel-convolutional neural network for infrared and visible image fusion [J]. IEEE Access, 2026, 14: 35396-35409.
[9] Tang W, He F, Liu Y. ITFuse: an interactive transformer for infrared and visible image fusion [J]. Pattern Recognition, 2024, 156: 110822-110834.
[10] 华怡坦, 黄影平, 过文昊. 基于CNN和Transformer点云图像融合的道路检测[J]. 应用科学学报, 2024, 42(4): 695-708. Hua Y T, Huang Y P, Guo W H. Fusion of point-cloud and image for road segmentation using CNN and Transformer [J]. Journal of Applied Sciences, 2024, 42(4): 695-708. (in Chinese)
[11] Hou Q, Lu C Z, Cheng M M, et al. Conv2Former: a simple transformer-style convnet for visual recognition [J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024, 46(12): 8274-8283.
[12] Li H, Yang Z, Zhang Y, et al. MulFS-CAP: multimodal fusion-supervised cross-modality alignment perception for unregistered infrared-visible image fusion [J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025, 47(5): 3673-3690.
[13] Yang Z, Zhang Y, Li H, et al. Instruction-driven fusion of Infrared-visible images: tailoring for diverse downstream tasks [J]. Information Fusion, 2025, 121: 103148-103162.
[14] Tang L, Yuan J, Ma J. Image fusion in the loop of high-level vision tasks: a semantic-aware real-time infrared and visible image fusion network [J]. Information Fusion, 2022, 82: 28-42.
[15] Bao W, Feng Z, Du Y, et al. A dual-modal semantic guidance and differential feature complementation fusion method for infrared and visible image [J]. Signal, Image and Video Processing, 2025, 19(2): 150-164.
[16] Yu C, Wang J, Peng C, et al. BiSeNet: bilateral segmentation network for real-time semantic segmentation [C] //European Conference on Computer Vision, 2018: 325-341.
[17] Xu H, Ma J, Jiang J, et al. U2Fusion: a unified unsupervised image fusion network [J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2020, 44(1): 502-518.
[18] Tang L, Yuan J, Zhang H, et al. PIAFusion: a progressive infrared and visible image fusion network based on illumination aware [J]. Information Fusion, 2022, 83: 79-92.
[19] Liu J, Fan X, Huang Z, et al. Target-aware dual adversarial learning and a multi-scenario multimodality benchmark to fuse infrared and visible for object detection [C] //IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022: 5802-5811.
[20] Zhang W, Sun H, Zhou B. TBRAFusion: infrared and visible image fusion based on two-branch residual attention transformer [J]. Electronic Research Archive, 2025, 33(1): 158-180.
[21] Zhao Z, Bai H, Zhang J, et al. CDDFuse: correlation-driven dual-branch feature decomposition for multi-modality image fusion [C] //IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023: 5906-5916.
[22] Li H, Wu X J. CrossFuse: a novel cross attention mechanism based infrared and visible image fusion approach [J]. Information Fusion, 2024, 103: 102147-102158.
[23] Xie X, Cui Y, Tan T, et al. FusionMamba: dynamic feature enhancement for multimodal image fusion with mamba [J]. Visual Intelligence, 2024, 2(1): 37-54.
[24] Tang L, Yan Q, Xiang X, et al. C2RF: bridging multi-modal image registration and fusion via commonality mining and contrastive learning [J]. International Journal of Computer Vision, 2025: 1-19.