智能视觉感知

上下文感知与模态对齐增强的语音翻译方法

  • 杨松涛 ,
  • 高盛祥 ,
  • 欧阳飞 ,
  • 余正涛
展开
  • 1. 昆明理工大学 信息工程与自动化学院, 云南 昆明 650500;
    2. 昆明理工大学 云南省人工智能重点实验室, 云南 昆明 650500

收稿日期: 2026-02-02

  网络出版日期: 2026-08-01

基金资助

国家自然科学基金(No.U24A20334,No.62376111);云南省科技重大专项计划项目(No.202303AP140008,No.202502AD080014,No.202502AD080014);云南省基础研究重大项目(No.202401BC070021)

Speech Translation Method Enhanced with Context Awareness and Modality Alignment

  • YANG Songtao ,
  • GAO Shengxiang ,
  • OUYANG Fei ,
  • YU Zhengtao
Expand
  • 1. Faculty of Information Engineering and Automation, Kunming University of Science and Technology, Kunming 650500, Yunnan, China;
    2. Key Laboratory of Artificial Intelligence in Yunnan Province, Kunming University of Science and Technology, Kunming 650500, Yunnan, China

Received date: 2026-02-02

  Online published: 2026-08-01

摘要

针对语音翻译任务中模态对齐模块对语音表征上下文信息感知不足的问题,提出了一种上下文语义增强型语音转文本的翻译方法。在语音表征基础上,引入基于局部窗口方差感知与相似度度量的上下文语义重组机制,对帧级语音表征进行语义感知的表征压缩,以获得上下文语义增强的语音最终表示。同时引入全局知识模态适配器,进一步实现语音上下文表征与文本的模态对齐。最终,将融合后的多模态表征输入至语音大模型解码器,由其解码生成目标语言的翻译文本。实验结果表明,该方法可有效提取包含上下文语义信息的语音表征,并在下游翻译任务中取得更优的性能。

本文引用格式

杨松涛 , 高盛祥 , 欧阳飞 , 余正涛 . 上下文感知与模态对齐增强的语音翻译方法[J]. 应用科学学报, 2026 , 44(4) : 571 -584 . DOI: 10.3969/j.issn.0255-8297.2026.04.004

Abstract

To address the problem of insufficient perception of contextual information in speech representations by modality alignment modules in speech translation tasks, a contextual semantic-enhanced speech-to-text translation method was proposed. Building upon speech representations, a contextual semantic reorganization mechanism based on local window variance perception and similarity measurement was introduced to perform semantic-aware representation compression on frame-level speech representations, so as to obtain final speech representations with contextual semantic enhancement. Meanwhile, a global knowledge modality adapter was introduced to further achieve modality alignment between the speech contextual representation and the text. Finally, the fused multimodal representation was input to a large speech model decoder, which decoded it to generate the translated text of the target language. Experimental results demonstrate that the method effectively extracts speech representations containing contextual semantic information and achieves better performance in downstream translation tasks.

参考文献

[1] Baevski A, Zhou Y, Mohamed A, et al. wav2vec 2.0:a framework for self-supervised learning of speech representations[PP/OL].V1.(2020-06-20)[2026-02-02]. https://arxiv.org/abs/2006.11477.
[2] Gulati A, Qin J, Chiu C C, et al. Conformer:convolution-augmented transformer for speech recognition[PP/OL].(2019-11-02)[2026-02-02]. https://aclanthology.org/2019.iwslt-1.29.pdf.
[3] Cheng Q, Fan M Y, Han Y Q, et al. Breaking the data barrier:towards robust speech translation via adversarial stability training[PP/OL].V1.(2019-09-25)[2026-02-02].https://arxiv.org/abs/1909.11430.
[4] Beck D, Cohn T, Haffari G. Neural speech translation using lattice transformations and graph networks[C] //13th Workshop on Graph-Based Methods for Natural Language Processing, 2019:26-31.
[5] Sperber M, Paulik M. Speech translation and the end-to-end promise:taking stock of where we are[C] //58th Annual Meeting of the Association for Computational Linguistics, 2020:7409-7421.
[6] Deshmukh S, Elizalde B, Singh R, et al. Pengi:an audio language model for audio tasks[PP/OL].V1.(2023-05-19)[2026-02-02]. https://arxiv.org/abs/2305.11834.
[7] Gong Y, Luo H, Liu A H, et al. Listen, think, and understand[PP/OL].V1.(2023-05-18)[2026-02-02]. https://arxiv.org/abs/2305.10790.
[8] Gong Y, Liu A H, Luo H Y, et al. Joint audio and speech understanding[C] //IEEE Automatic Speech Recognition and Understanding Workshop, 2023:1-8.
[9] Tang C, Yu W, Sun G, et al. Salmonn:towards generic hearing abilities for large language models[PP/OL].V2.(2024-04-08)[2026-02-02]. https://arxiv.org/pdf/2310.132892024.
[10] Ghosh S, Kumar S, Seth A, et al. GAMA:a large audio-language model with advanced audio understanding and complex reasoning abilities[C] //Conference on Empirical Methods in Natural Language Processing, 2024:6288-6313.
[11] Huang R J, Li M Z, Yang D C, et al. AudioGPT:understanding and generating speech, music,sound, and talking head[C] //AAAI Conference on Artificial Intelligence, 2024, 38(21):23802-23804.
[12] Zhang D, Li S M, Zhang X, et al. SpeechGPT:empowering large language models with intrinsic cross-modal conversational abilities[C] //Conference on Empirical Methods in Natural Language Processing, 2023:15757-15773.
[13] Wang T R, Zhou L, Zhang Z Q, et al. VioLA:conditional language models for speech recognition, synthesis, and translation[J]. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2024, 32:3709-3716.
[14] Chen X, Zhang S Y, Bai Q B, et al. LLaST:improved end-to-end speech translation system leveraged by large language models[C] //Conference on Empirical Methods in Natural Language Processing, 2024:6976-6987.
[15] Zhang Y H, Xu C, Li B, et al. Rethinking and improving multi-task learning for end-to-end speech translation[C] //Conference on Empirical Methods in Natural Language Processing, 2023:10753-10765.
[16] Ni J J, Ma Y K, Wang W, et al. Adaptive knowledge distillation between text and speech pre-trained models[PP/OL].V1.(2023-03-07)[2026-02-02]. https://arxiv.org/abs/2303.03600.
[17] Du Y X, Pan Y C, Ma Z Y, et al. Making LLMs better many-to-many speech-to-text translators with curriculum learning[C] //The 63rd Annual Meeting of the Association for Computational Linguistics, 2025:12466-12478.
[18] Liu H L, Chen A D, Chen K H, et al. Adaptive inner speech text alignment for LLM-based speech translation[M] //Natural Language Processing and Chinese Computing. Singapore 2025.
[19] Wang C H, Wu A, Gu J T, et al. CoVoST 2 and massively multilingual speech translation[PP/OL].V1.(2020-07-20)[2026-02-02]. https://arxiv.org/abs/2007.10310.
[20] Di Gangi M A, Cattoni R, Bentivogli L, et al. MuST-C:a multilingual speech translation corpus[C] //Conference of the North American Chapter of the Association for Computational Linguistics:Human Language Technologies, 2019:2012-2017.
[21] Radford A, Kim J W, Xu T, et al. Robust speech recognition via large-scale weak supervision[C] //International Conference on Machine Learning, 2023:28492-28518.
[22] Hu E J, Shen Y, Wallis P, et al. LoRA:low-rank adaptation of large language models[PP/OL].V1.(2021-06-17)[2026-02-02]. https://arxiv.org/abs/2106.09685.
[23] Post M. A call for clarity in reporting BLEU scores[PP/OL].[2026-02-02]. https://aclanthology.org/W18-6319/.
[24] Rei R, de Souza J G C, Alves D, et al. COMET-22:unbabel-IST 2022 submission for the metrics shared task[C] //7th Conference on Machine Translation, 2022:578-585.
文章导航

/