应用科学学报 ›› 2026, Vol. 44 ›› Issue (4): 571-584.doi: 10.3969/j.issn.0255-8297.2026.04.004

• 智能视觉感知 • 上一篇    下一篇

上下文感知与模态对齐增强的语音翻译方法

杨松涛1,2, 高盛祥1,2, 欧阳飞1,2, 余正涛1,2   

  1. 1. 昆明理工大学 信息工程与自动化学院, 云南 昆明 650500;
    2. 昆明理工大学 云南省人工智能重点实验室, 云南 昆明 650500
  • 收稿日期:2026-02-02 发布日期:2026-08-01
  • 通信作者: 高盛祥,教授,博士生导师,研究方向为人工智能、自然语言处理、机器翻译、信息检索、语音识别与合成。E-mail:gaoshengxiang.yn@foxmail.com E-mail:gaoshengxiang.yn@foxmail.com
  • 基金资助:
    国家自然科学基金(No.U24A20334,No.62376111);云南省科技重大专项计划项目(No.202303AP140008,No.202502AD080014,No.202502AD080014);云南省基础研究重大项目(No.202401BC070021)

Speech Translation Method Enhanced with Context Awareness and Modality Alignment

YANG Songtao1,2, GAO Shengxiang1,2, OUYANG Fei1,2, YU Zhengtao1,2   

  1. 1. Faculty of Information Engineering and Automation, Kunming University of Science and Technology, Kunming 650500, Yunnan, China;
    2. Key Laboratory of Artificial Intelligence in Yunnan Province, Kunming University of Science and Technology, Kunming 650500, Yunnan, China
  • Received:2026-02-02 Published:2026-08-01

摘要: 针对语音翻译任务中模态对齐模块对语音表征上下文信息感知不足的问题,提出了一种上下文语义增强型语音转文本的翻译方法。在语音表征基础上,引入基于局部窗口方差感知与相似度度量的上下文语义重组机制,对帧级语音表征进行语义感知的表征压缩,以获得上下文语义增强的语音最终表示。同时引入全局知识模态适配器,进一步实现语音上下文表征与文本的模态对齐。最终,将融合后的多模态表征输入至语音大模型解码器,由其解码生成目标语言的翻译文本。实验结果表明,该方法可有效提取包含上下文语义信息的语音表征,并在下游翻译任务中取得更优的性能。

关键词: 语音翻译, 大语言模型, 语义增强, 模态对齐

Abstract: To address the problem of insufficient perception of contextual information in speech representations by modality alignment modules in speech translation tasks, a contextual semantic-enhanced speech-to-text translation method was proposed. Building upon speech representations, a contextual semantic reorganization mechanism based on local window variance perception and similarity measurement was introduced to perform semantic-aware representation compression on frame-level speech representations, so as to obtain final speech representations with contextual semantic enhancement. Meanwhile, a global knowledge modality adapter was introduced to further achieve modality alignment between the speech contextual representation and the text. Finally, the fused multimodal representation was input to a large speech model decoder, which decoded it to generate the translated text of the target language. Experimental results demonstrate that the method effectively extracts speech representations containing contextual semantic information and achieves better performance in downstream translation tasks.

Key words: speech translation, large language model, semantic enhancement, modality alignment

中图分类号: