To address the problem of insufficient perception of contextual information in speech representations by modality alignment modules in speech translation tasks, a contextual semantic-enhanced speech-to-text translation method was proposed. Building upon speech representations, a contextual semantic reorganization mechanism based on local window variance perception and similarity measurement was introduced to perform semantic-aware representation compression on frame-level speech representations, so as to obtain final speech representations with contextual semantic enhancement. Meanwhile, a global knowledge modality adapter was introduced to further achieve modality alignment between the speech contextual representation and the text. Finally, the fused multimodal representation was input to a large speech model decoder, which decoded it to generate the translated text of the target language. Experimental results demonstrate that the method effectively extracts speech representations containing contextual semantic information and achieves better performance in downstream translation tasks.
YANG Songtao
,
GAO Shengxiang
,
OUYANG Fei
,
YU Zhengtao
. Speech Translation Method Enhanced with Context Awareness and Modality Alignment[J]. Journal of Applied Sciences, 2026
, 44(4)
: 571
-584
.
DOI: 10.3969/j.issn.0255-8297.2026.04.004
[1] Baevski A, Zhou Y, Mohamed A, et al. wav2vec 2.0:a framework for self-supervised learning of speech representations[PP/OL].V1.(2020-06-20)[2026-02-02]. https://arxiv.org/abs/2006.11477.
[2] Gulati A, Qin J, Chiu C C, et al. Conformer:convolution-augmented transformer for speech recognition[PP/OL].(2019-11-02)[2026-02-02]. https://aclanthology.org/2019.iwslt-1.29.pdf.
[3] Cheng Q, Fan M Y, Han Y Q, et al. Breaking the data barrier:towards robust speech translation via adversarial stability training[PP/OL].V1.(2019-09-25)[2026-02-02].https://arxiv.org/abs/1909.11430.
[4] Beck D, Cohn T, Haffari G. Neural speech translation using lattice transformations and graph networks[C] //13th Workshop on Graph-Based Methods for Natural Language Processing, 2019:26-31.
[5] Sperber M, Paulik M. Speech translation and the end-to-end promise:taking stock of where we are[C] //58th Annual Meeting of the Association for Computational Linguistics, 2020:7409-7421.
[6] Deshmukh S, Elizalde B, Singh R, et al. Pengi:an audio language model for audio tasks[PP/OL].V1.(2023-05-19)[2026-02-02]. https://arxiv.org/abs/2305.11834.
[7] Gong Y, Luo H, Liu A H, et al. Listen, think, and understand[PP/OL].V1.(2023-05-18)[2026-02-02]. https://arxiv.org/abs/2305.10790.
[8] Gong Y, Liu A H, Luo H Y, et al. Joint audio and speech understanding[C] //IEEE Automatic Speech Recognition and Understanding Workshop, 2023:1-8.
[9] Tang C, Yu W, Sun G, et al. Salmonn:towards generic hearing abilities for large language models[PP/OL].V2.(2024-04-08)[2026-02-02]. https://arxiv.org/pdf/2310.132892024.
[10] Ghosh S, Kumar S, Seth A, et al. GAMA:a large audio-language model with advanced audio understanding and complex reasoning abilities[C] //Conference on Empirical Methods in Natural Language Processing, 2024:6288-6313.
[11] Huang R J, Li M Z, Yang D C, et al. AudioGPT:understanding and generating speech, music,sound, and talking head[C] //AAAI Conference on Artificial Intelligence, 2024, 38(21):23802-23804.
[12] Zhang D, Li S M, Zhang X, et al. SpeechGPT:empowering large language models with intrinsic cross-modal conversational abilities[C] //Conference on Empirical Methods in Natural Language Processing, 2023:15757-15773.
[13] Wang T R, Zhou L, Zhang Z Q, et al. VioLA:conditional language models for speech recognition, synthesis, and translation[J]. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2024, 32:3709-3716.
[14] Chen X, Zhang S Y, Bai Q B, et al. LLaST:improved end-to-end speech translation system leveraged by large language models[C] //Conference on Empirical Methods in Natural Language Processing, 2024:6976-6987.
[15] Zhang Y H, Xu C, Li B, et al. Rethinking and improving multi-task learning for end-to-end speech translation[C] //Conference on Empirical Methods in Natural Language Processing, 2023:10753-10765.
[16] Ni J J, Ma Y K, Wang W, et al. Adaptive knowledge distillation between text and speech pre-trained models[PP/OL].V1.(2023-03-07)[2026-02-02]. https://arxiv.org/abs/2303.03600.
[17] Du Y X, Pan Y C, Ma Z Y, et al. Making LLMs better many-to-many speech-to-text translators with curriculum learning[C] //The 63rd Annual Meeting of the Association for Computational Linguistics, 2025:12466-12478.
[18] Liu H L, Chen A D, Chen K H, et al. Adaptive inner speech text alignment for LLM-based speech translation[M] //Natural Language Processing and Chinese Computing. Singapore 2025.
[19] Wang C H, Wu A, Gu J T, et al. CoVoST 2 and massively multilingual speech translation[PP/OL].V1.(2020-07-20)[2026-02-02]. https://arxiv.org/abs/2007.10310.
[20] Di Gangi M A, Cattoni R, Bentivogli L, et al. MuST-C:a multilingual speech translation corpus[C] //Conference of the North American Chapter of the Association for Computational Linguistics:Human Language Technologies, 2019:2012-2017.
[21] Radford A, Kim J W, Xu T, et al. Robust speech recognition via large-scale weak supervision[C] //International Conference on Machine Learning, 2023:28492-28518.
[22] Hu E J, Shen Y, Wallis P, et al. LoRA:low-rank adaptation of large language models[PP/OL].V1.(2021-06-17)[2026-02-02]. https://arxiv.org/abs/2106.09685.
[23] Post M. A call for clarity in reporting BLEU scores[PP/OL].[2026-02-02]. https://aclanthology.org/W18-6319/.
[24] Rei R, de Souza J G C, Alves D, et al. COMET-22:unbabel-IST 2022 submission for the metrics shared task[C] //7th Conference on Machine Translation, 2022:578-585.