Journal of Applied Sciences ›› 2026, Vol. 44 ›› Issue (4): 571-584.doi: 10.3969/j.issn.0255-8297.2026.04.004

• Intelligent Visual Perception • Previous Articles     Next Articles

Speech Translation Method Enhanced with Context Awareness and Modality Alignment

YANG Songtao1,2, GAO Shengxiang1,2, OUYANG Fei1,2, YU Zhengtao1,2   

  1. 1. Faculty of Information Engineering and Automation, Kunming University of Science and Technology, Kunming 650500, Yunnan, China;
    2. Key Laboratory of Artificial Intelligence in Yunnan Province, Kunming University of Science and Technology, Kunming 650500, Yunnan, China
  • Received:2026-02-02 Published:2026-08-01

Abstract: To address the problem of insufficient perception of contextual information in speech representations by modality alignment modules in speech translation tasks, a contextual semantic-enhanced speech-to-text translation method was proposed. Building upon speech representations, a contextual semantic reorganization mechanism based on local window variance perception and similarity measurement was introduced to perform semantic-aware representation compression on frame-level speech representations, so as to obtain final speech representations with contextual semantic enhancement. Meanwhile, a global knowledge modality adapter was introduced to further achieve modality alignment between the speech contextual representation and the text. Finally, the fused multimodal representation was input to a large speech model decoder, which decoded it to generate the translated text of the target language. Experimental results demonstrate that the method effectively extracts speech representations containing contextual semantic information and achieves better performance in downstream translation tasks.

Key words: speech translation, large language model, semantic enhancement, modality alignment

CLC Number: