Speech Translation with Speech Foundation Models and Large Language Models: What is There and What is Missing?
Marco Gaido, Sara Papi, Matteo Negri, Luisa Bentivogli
2024Year
6Citations
6Top-tier citations
Abstract
COSMIC TEDLIUM3 ASR, SQA TEDLIUM3, FLEURS en→es, fr, de, zh SLM Alpaca, CoVoST2, YouTube (in-house) ASR, ST, SIT No No CoVoST2
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 16c1b3eb-4399-44e7-a569-ff89bcf35591Cited by top-tier papers6
- MCIF: Multimodal Crosslingual Instruction-Following Benchmark from Scientific TalksSara Papi, Maike Züfle, Marco Gaido, Beatrice Savoldi et al.ICLR 2026 · 20 citations
- Towards Building Large Scale Datasets and State-of-the-Art Automatic Speech Translation Systems for 14 Indian LanguagesAshwin Sankar, Sparsh Jain, Nikhil Narasimhan, Devilal Choudhary et al.ACL 2025 · 5 citations
- Summarizing Speech: A Comprehensive SurveyFabian Retkowski, Maike Züfle, Andreas Sudmann, Dinah Pfau et al.EMNLP 2025 · 3 citations
- SpeechQE: Estimating the Quality of Direct Speech TranslationHyoJung Han, Kevin Duh, Marine CarpuatEMNLP 2024 · 1 citation
- PLaST: Towards Paralinguistic-aware Speech TranslationYi Li, Rui Zhao, Ruiquan Zhang, Jinsong Su et al.AAAI 2026
Builds on23
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- LLM-Pruner: On the Structural Pruning of Large Language ModelsXinyin Ma, Gongfan Fang, Xinchao WangNeurIPS 2023 · 994 citations
- SALMONN: Towards Generic Hearing Abilities for Large Language ModelsChangli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen et al.ICLR 2024 · 557 citations
Related papers
- Align3R: Aligned Monocular Depth Estimation for Dynamic VideosJiahao Lu, Tianyu Huang, Peng Li, Zhiyang Dou et al.CVPR 2025
- SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and TrainingJierun Chen, Dongting Hu, Xijie Huang, Huseyin Coskun et al.CVPR 2025
- Learning Visual Generative Priors without TextShuailei Ma, Kecheng Zheng, Ying Wei, Wei Wu et al.CVPR 2025
- Guided Score identity Distillation for Data-Free One-Step Text-to-Image GenerationMingyuan Zhou, Zhendong Wang, Huangjie Zheng, Hai HuangICLR 2025
- Hash3D: Training-free Acceleration for 3D GenerationXingyi Yang, Songhua Liu, Xinchao WangCVPR 2025
