Artiflcial Intelligence Technology and Applications

Multi-modal Rumor Detection Method Fusing Image and Text Features

  • GAO Guangliang ,
  • LIANG Weichao ,
  • ZHU Tao ,
  • HONG Lei ,
  • XIA Lingling
Expand
  • 1. School of National Security, Jiangsu Police Institute, Nanjing 210031, Jiangsu, China;
    2. School of Computer Science and Engineering, Guangxi Normal University, Guilin 541004, Guangxi, China;
    3. School of Digital and Intelligent Police Technology, Jiangsu Police Institute, Nanjing 210031, Jiangsu, China;
    4. School of Cyberspace Security, Jiangsu Police Institute, Nanjing 210031, Jiangsu, China

Received date: 2025-12-17

  Online published: 2026-06-23

Abstract

To further improve the effectiveness and stability of rumor detection, a multimodal rumor detection method was proposed that integrated image and text features with adaptive noise resistance and semantic restoration. First, a three-stage processing pipeline consisting of pretrained encoding, dynamic pooling, and multi-head enhancement was designed to encode rumor-related texts and comments into semantic vectors. Then,two parallel modules were constructed: one for visual feature extraction and the other for optical character feature extraction. These modules encoded rumor images and explicit text within them into complementary enhanced image vectors. Finally, adaptive gating and dynamic cross-attention mechanisms were used to filter noise and enhance semantics,achieving local alignment and global integration of image and text information. Experimental results show that, compared with baseline algorithms, the proposed method can effectively capture deep correlations between image and text information and improve credibility and practicality of rumor detection results.

Cite this article

GAO Guangliang , LIANG Weichao , ZHU Tao , HONG Lei , XIA Lingling . Multi-modal Rumor Detection Method Fusing Image and Text Features[J]. Journal of Applied Sciences, 2026 , 44(3) : 503 -514 . DOI: 10.3969/j.issn.0255-8297.2026.03.011

References

[1] Chen Y X, Li D S, Zhang P, et al. Cross-modal ambiguity learning for multimodal fake news detection [C]//ACM Web Conference, 2022: 2897-2905.
[2] Jing J, Wu H C, Sun J, et al. Multimodal fake news detection via progressive fusion networks [J]. Information Processing and Management, 2023, 60(1): 103120-103133.
[3] Singhal S, Pandey T, Mrig S, et al. Leveraging intra and inter modality relationship for multimodal fake news detection [C]//Companion Web Conference, 2022: 726-734.
[4] 蒋超, 朱学芳. 基于模态融合增强的谣言检测研究[J]. 数据分析与知识发现, 2025, 9(7): 26-37. Jiang C, Zhu X F. Research on rumour detection based on modal fusion enhancement [J]. Data Analysis and Knowledge Discovery, 2025, 9(7): 26-37. (in Chinese)
[5] Wang L Z, Zhang C, Xu H B, et al. Cross-modal contrastive learning for multimodal fake news detection [C]//31st International Conference on Multimedia, 2023: 5696-5704.
[6] Meel P, Vishwakarma D K. Han, image captioning, and forensics ensemble multimodal fake news detection [J]. Information Sciences, 2021, 567: 23-41.
[7] Song C G, Ning N W, Zhang Y L, et al. A multimodal fake news detection model based on crossmodal attention residual and multichannel convolutional neural networks [J]. Information Processing and Management, 2021, 58(1): 102437-102445.
[8] Kumari R, Ekbal A. AMFB: attention based multimodal factorized bilinear pooling for multimodal fake news detection [J]. Expert Systems with Applications, 2021, 184(1): 115412-115424.
[9] Segura-Bedmar I, Alonso-Bartolome S. Multimodal fake news detection [J]. Information, 2022, 13(6): 284-300.
[10] Hua J H, Cui X D, Li X H, et al. Multimodal fake news detection through data augmentationbased contrastive learning [J]. Applied Soft Computing, 2023, 136: 110125-110134.
[11] Giachanou A, Zhang G B, Rosso P. Multimodal fake news detection with textual, visual and semantic information [C]//Text, Speech, and Dialogue, 2020: 30-38.
[12] Liu F, Xu G H, Wu Q, et al. Cascade reasoning network for text-based visual question answering [C]//28th International Conference on Multimedia, 2020: 4060-4069.
[13] Han W, Huang H T, Han T. Finding the evidence: localization-aware answer prediction for text visual question answering [C]//28th International Conference on Computational Linguistics, 2020: 3118-3131.
[14] Wang C Y, Zhang M, Shi F, et al. A hybrid multimodal data fusion-based method for identifying gambling websites [J]. Electronics, 2022, 11(16): 1-18.
[15] Lin Z H, Xie J Y, Li Q. Multi-modal news event detection with external knowledge [J]. Information Processing and Management, 2024, 61(3): 103697-103714.
[16] 刘先博, 向澳, 杜彦辉. 融合图文多粒度情感特征的多模态谣言检测方法[J]. 计算机科学与探索, 2025, 19(4): 1021-1035. Liu X B, Xiang A, Du Y H. Multimodal rumor detection method based on multi-granularity emotional features of image-text [J]. Journal of Frontiers of Computer Science and Technology, 2025, 19(4): 1021-1035. (in Chinese)
[17] Zhang L W B, Wilson R, Sumner M, et al. Advanced multimodal fusion method for very short-term solar irradiance forecasting using sky images and meteorological data: a gate and transformer mechanism approach [J]. Renewable Energy, 2023, 216: 118952-118959.
[18] 黄学坚, 马廷淮, 荣欢, 等. 融合外部知识与证据的场景图注意力网络多模态谣言检测[J]. 计算机学报, 2025, 48(9): 2159-2180. Huang X J, Ma T H, Rong H, et al. Multi-modal rumor detection with scene graph attention networks integrating external knowledge and evidence [J]. Chinese Journal of Computers, 2025, 48(9): 2159-2180. (in Chinese)
[19] Kumar A, Vepa J. Gated mechanism for attention based multimodal sentiment analysis [C]//45th IEEE International Conference on Acoustics, Speech and Signal Processing, 2020: 4477-4481.
[20] Jin Z W, Cao J, Guo H, et al. Multimodal fusion with recurrent neural networks for rumor detection on microblogs [C]//25th ACM International Conference on Multimedia, 2017: 795-816.
[21] Wu Y, Zhan P W, Zhang Y J, et al. Multimodal fusion with co-attention networks for fake news detection [C]//Findings of the Association for Computational Linguistics, 2021: 2560-2569.
[22] Peng L W, Jian S L, Li D S, et al. MRML: multimodal rumor detection by deep metric learning [C]//48th IEEE International Conference on Acoustics, Speech and Signal Processing, 2023: 10096188-10096193.
[23] Xu F, Zeng L, Huang Q, et al. Hierarchical graph attention networks for multi-modal rumor detection on social media [J]. Neurocomputing, 2024, 569: 127112-127123.
Outlines

/