[1] Baevski A, Zhou Y, Mohamed A, et al. wav2vec 2.0:a framework for self-supervised learning of speech representations[PP/OL].V1.(2020-06-20)[2026-02-02]. https://arxiv.org/abs/2006.11477. [2] Gulati A, Qin J, Chiu C C, et al. Conformer:convolution-augmented transformer for speech recognition[PP/OL].(2019-11-02)[2026-02-02]. https://aclanthology.org/2019.iwslt-1.29.pdf. [3] Cheng Q, Fan M Y, Han Y Q, et al. Breaking the data barrier:towards robust speech translation via adversarial stability training[PP/OL].V1.(2019-09-25)[2026-02-02].https://arxiv.org/abs/1909.11430. [4] Beck D, Cohn T, Haffari G. Neural speech translation using lattice transformations and graph networks[C] //13th Workshop on Graph-Based Methods for Natural Language Processing, 2019:26-31. [5] Sperber M, Paulik M. Speech translation and the end-to-end promise:taking stock of where we are[C] //58th Annual Meeting of the Association for Computational Linguistics, 2020:7409-7421. [6] Deshmukh S, Elizalde B, Singh R, et al. Pengi:an audio language model for audio tasks[PP/OL].V1.(2023-05-19)[2026-02-02]. https://arxiv.org/abs/2305.11834. [7] Gong Y, Luo H, Liu A H, et al. Listen, think, and understand[PP/OL].V1.(2023-05-18)[2026-02-02]. https://arxiv.org/abs/2305.10790. [8] Gong Y, Liu A H, Luo H Y, et al. Joint audio and speech understanding[C] //IEEE Automatic Speech Recognition and Understanding Workshop, 2023:1-8. [9] Tang C, Yu W, Sun G, et al. Salmonn:towards generic hearing abilities for large language models[PP/OL].V2.(2024-04-08)[2026-02-02]. https://arxiv.org/pdf/2310.132892024. [10] Ghosh S, Kumar S, Seth A, et al. GAMA:a large audio-language model with advanced audio understanding and complex reasoning abilities[C] //Conference on Empirical Methods in Natural Language Processing, 2024:6288-6313. [11] Huang R J, Li M Z, Yang D C, et al. AudioGPT:understanding and generating speech, music,sound, and talking head[C] //AAAI Conference on Artificial Intelligence, 2024, 38(21):23802-23804. [12] Zhang D, Li S M, Zhang X, et al. SpeechGPT:empowering large language models with intrinsic cross-modal conversational abilities[C] //Conference on Empirical Methods in Natural Language Processing, 2023:15757-15773. [13] Wang T R, Zhou L, Zhang Z Q, et al. VioLA:conditional language models for speech recognition, synthesis, and translation[J]. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2024, 32:3709-3716. [14] Chen X, Zhang S Y, Bai Q B, et al. LLaST:improved end-to-end speech translation system leveraged by large language models[C] //Conference on Empirical Methods in Natural Language Processing, 2024:6976-6987. [15] Zhang Y H, Xu C, Li B, et al. Rethinking and improving multi-task learning for end-to-end speech translation[C] //Conference on Empirical Methods in Natural Language Processing, 2023:10753-10765. [16] Ni J J, Ma Y K, Wang W, et al. Adaptive knowledge distillation between text and speech pre-trained models[PP/OL].V1.(2023-03-07)[2026-02-02]. https://arxiv.org/abs/2303.03600. [17] Du Y X, Pan Y C, Ma Z Y, et al. Making LLMs better many-to-many speech-to-text translators with curriculum learning[C] //The 63rd Annual Meeting of the Association for Computational Linguistics, 2025:12466-12478. [18] Liu H L, Chen A D, Chen K H, et al. Adaptive inner speech text alignment for LLM-based speech translation[M] //Natural Language Processing and Chinese Computing. Singapore 2025. [19] Wang C H, Wu A, Gu J T, et al. CoVoST 2 and massively multilingual speech translation[PP/OL].V1.(2020-07-20)[2026-02-02]. https://arxiv.org/abs/2007.10310. [20] Di Gangi M A, Cattoni R, Bentivogli L, et al. MuST-C:a multilingual speech translation corpus[C] //Conference of the North American Chapter of the Association for Computational Linguistics:Human Language Technologies, 2019:2012-2017. [21] Radford A, Kim J W, Xu T, et al. Robust speech recognition via large-scale weak supervision[C] //International Conference on Machine Learning, 2023:28492-28518. [22] Hu E J, Shen Y, Wallis P, et al. LoRA:low-rank adaptation of large language models[PP/OL].V1.(2021-06-17)[2026-02-02]. https://arxiv.org/abs/2106.09685. [23] Post M. A call for clarity in reporting BLEU scores[PP/OL].[2026-02-02]. https://aclanthology.org/W18-6319/. [24] Rei R, de Souza J G C, Alves D, et al. COMET-22:unbabel-IST 2022 submission for the metrics shared task[C] //7th Conference on Machine Translation, 2022:578-585. |