Intelligent Visual Perception

Semantic-Driven Real-Time Infrared and Visible Image Fusion Based on Spatial-Frequency Perception and Differential Complementation

  • ZHANG Zujun ,
  • GAN Rui ,
  • LI Wei ,
  • CHEN Jiujiu ,
  • XIE Yu ,
  • XIONG Bangshu
Expand
  • 1. School of Information Engineering, Nanchang Hangkong University, Nanchang 330063, Jiangxi, China;
    2. Jiangxi Communications Investment Group Co., Ltd., Nanchang 330013, Jiangxi, China

Received date: 2026-02-05

  Online published: 2026-08-01

Abstract

To address the limitations of existing methods, such as limited global feature extraction, insufficient fusion of cross-modal complementary features, and the neglect of downstream visual task requirements, a semantic-driven real-time infrared and visible image fusion method based on spatial-frequency perception and differential complementation is proposed. First, a spatial-frequency dual-branch perception module was designed to effectively capture the local spatial details and global dependencies of the image. Second, a cross-modality differential feature complementation module was designed to fully integrate the advantageous features of different modalities. Furthermore, a semantic-driven joint training framework was constructed, using semantic segmentation loss to guide the fusion network to retain more semantic information, thereby improving the performance of downstream advanced visual tasks. Experimental results on the public datasets RoadScene, MSRS, and M3FD show that, compared with mainstream methods, the mutual information and visual fidelity of the fused images are improved by an average of 3.20% and 3.79%, respectively, by the proposed method; in the semantic segmentation task, the mean intersection over union is improved by 2.98%; in the object detection task, the mean average precision is improved by 5.64%. In terms of running efficiency, the processing frame rate reaches 38.99 FPS, meeting the real-time requirements of engineering applications.

Cite this article

ZHANG Zujun , GAN Rui , LI Wei , CHEN Jiujiu , XIE Yu , XIONG Bangshu . Semantic-Driven Real-Time Infrared and Visible Image Fusion Based on Spatial-Frequency Perception and Differential Complementation[J]. Journal of Applied Sciences, 2026 , 44(4) : 685 -700 . DOI: 10.3969/j.issn.0255-8297.2026.04.012

References

[1] 金安安, 李祥, 张丽, 等. 基于NSCT与压缩感知的红外影像融合[J]. 应用科学学报, 2022, 40(1): 80-92. Jin A A, Li X, Zhang L, et al. Infrared image fusion based on NSCT and compressed sensing [J]. Journal of Applied Sciences, 2022, 40(1): 80-92. (in Chinese)
[2] Li Y, Li D, Fang S, et al. Multimodal image fusion network with prior-guided dynamic degradation removal for extreme environment perception [J]. Scientific Reports, 2025, 15(1): 40530- 40551.
[3] Li S, Yang B, Hu J. Performance comparison of different multi-resolution transforms for image fusion [J]. Information Fusion, 2011, 12(2): 74-84.
[4] Li H, Wu X J. Multi-focus image fusion using dictionary learning and low-rank representation [C] //International Conference on Image and Graphics, 2017: 675-686.
[5] Li S, Yin H, Fang L. Group-sparse representation with dictionary learning for medical image denoising and fusion [J]. IEEE Transactions on Biomedical Engineering, 2012, 59(12): 3450-3459.
[6] Ma J, Tang L, Xu M, et al. STDFusionNet: an infrared and visible image fusion network based on salient target detection [J]. IEEE Transactions on Instrumentation and Measurement, 2021, 70: 1-13.
[7] Xu J, Zhang J. Infrared and visible image fusion based on Haar wavelet downsampling and multi-scale feature aggregation [J]. Signal, Image and Video Processing, 2025, 19(14): 1200- 1208.
[8] Duan X, Liu Y, Yang Z, et al. CCFusion: a novel channel-convolutional neural network for infrared and visible image fusion [J]. IEEE Access, 2026, 14: 35396-35409.
[9] Tang W, He F, Liu Y. ITFuse: an interactive transformer for infrared and visible image fusion [J]. Pattern Recognition, 2024, 156: 110822-110834.
[10] 华怡坦, 黄影平, 过文昊. 基于CNN和Transformer点云图像融合的道路检测[J]. 应用科学学报, 2024, 42(4): 695-708. Hua Y T, Huang Y P, Guo W H. Fusion of point-cloud and image for road segmentation using CNN and Transformer [J]. Journal of Applied Sciences, 2024, 42(4): 695-708. (in Chinese)
[11] Hou Q, Lu C Z, Cheng M M, et al. Conv2Former: a simple transformer-style convnet for visual recognition [J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024, 46(12): 8274-8283.
[12] Li H, Yang Z, Zhang Y, et al. MulFS-CAP: multimodal fusion-supervised cross-modality alignment perception for unregistered infrared-visible image fusion [J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025, 47(5): 3673-3690.
[13] Yang Z, Zhang Y, Li H, et al. Instruction-driven fusion of Infrared-visible images: tailoring for diverse downstream tasks [J]. Information Fusion, 2025, 121: 103148-103162.
[14] Tang L, Yuan J, Ma J. Image fusion in the loop of high-level vision tasks: a semantic-aware real-time infrared and visible image fusion network [J]. Information Fusion, 2022, 82: 28-42.
[15] Bao W, Feng Z, Du Y, et al. A dual-modal semantic guidance and differential feature complementation fusion method for infrared and visible image [J]. Signal, Image and Video Processing, 2025, 19(2): 150-164.
[16] Yu C, Wang J, Peng C, et al. BiSeNet: bilateral segmentation network for real-time semantic segmentation [C] //European Conference on Computer Vision, 2018: 325-341.
[17] Xu H, Ma J, Jiang J, et al. U2Fusion: a unified unsupervised image fusion network [J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2020, 44(1): 502-518.
[18] Tang L, Yuan J, Zhang H, et al. PIAFusion: a progressive infrared and visible image fusion network based on illumination aware [J]. Information Fusion, 2022, 83: 79-92.
[19] Liu J, Fan X, Huang Z, et al. Target-aware dual adversarial learning and a multi-scenario multimodality benchmark to fuse infrared and visible for object detection [C] //IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022: 5802-5811.
[20] Zhang W, Sun H, Zhou B. TBRAFusion: infrared and visible image fusion based on two-branch residual attention transformer [J]. Electronic Research Archive, 2025, 33(1): 158-180.
[21] Zhao Z, Bai H, Zhang J, et al. CDDFuse: correlation-driven dual-branch feature decomposition for multi-modality image fusion [C] //IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023: 5906-5916.
[22] Li H, Wu X J. CrossFuse: a novel cross attention mechanism based infrared and visible image fusion approach [J]. Information Fusion, 2024, 103: 102147-102158.
[23] Xie X, Cui Y, Tan T, et al. FusionMamba: dynamic feature enhancement for multimodal image fusion with mamba [J]. Visual Intelligence, 2024, 2(1): 37-54.
[24] Tang L, Yan Q, Xiang X, et al. C2RF: bridging multi-modal image registration and fusion via commonality mining and contrastive learning [J]. International Journal of Computer Vision, 2025: 1-19.
Outlines

/