SIFT-50M: A Large-Scale Multilingual Dataset for Speech Instruction Fine-Tuning
Prabhat Pandey, Rupak Vignesh Swaminathan, K. V. Vijay Girish, Arunasish Sen, Jian Xie, Grant P. Strimel, Andreas Schwarz
Abstract
We introduce SIFT (Speech Instruction Fine-Tuning), a 50M-example dataset designed for instruction fine-tuning and pre-training of speech-text large language models (LLMs). SIFT-50M is built from publicly available speech corpora, which collectively contain 14K hours of speech, and leverages LLMs along with off-the-shelf expert models. The dataset spans five languages, encompassing a diverse range of speech understanding as well as controllable speech generation instructions. Using SIFT-50M, we train SIFT-LLM, which outperforms existing speech-text LLMs on instruction-following benchmarks while achieving competitive performance on foundational speech tasks. To support further research, we also introduce EvalSIFT, a benchmark dataset specifically designed to evaluate the instruction-following capabilities of speech-text LLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext daaeb27a-984f-47a8-aee3-5474338fef56Cited by top-tier papers4
- Music Flamingo: Scaling Music Understanding in Audio Language ModelsSreyan Ghosh, Arushi Goel, Lasha Koroshinadze, Sang-gil Lee et al.ICLR 2026 · 33 citations
- MCIF: Multimodal Crosslingual Instruction-Following Benchmark from Scientific TalksSara Papi, Maike Züfle, Marco Gaido, Beatrice Savoldi et al.ICLR 2026 · 20 citations
- Towards Holistic Evaluation of Large Audio-Language Models: A Comprehensive SurveyChih-Kai Yang, Neo S. Ho, Hung-yi LeeEMNLP 2025 · 7 citations
- Unlocking Speech–Text Compositional Powers: Instruction-Following Speech Language Models without Instruction TuningCongrui Du, Yang Zhang, Kaizhi Qian, Shiyu ChangICML 2026
Builds on4
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- SALMONN: Towards Generic Hearing Abilities for Large Language ModelsChangli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen et al.ICLR 2024 · 557 citations
- BigVGAN: A Universal Neural Vocoder with Large-Scale TrainingSang-gil Lee, Wei Ping, Boris Ginsburg, Bryan Catanzaro et al.ICLR 2023 · 46 citations
Related papers
- InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-trainingDingdong Wang, Jin Xu, Ruihang Chu, Zhifang Guo et al.ACL 2025 · 9 citations
- MaXIFE: Multilingual and Cross-lingual Instruction Following EvaluationYile Liu, Ziwei Ma, Xiu Jiang, Jinglu Hu et al.ACL 2025 · 5 citations
- SciRIFF: A Resource to Enhance Language Model Instruction-Following over Scientific LiteratureDavid Wadden, Kejian Shi, Jacob Morrison, Alan Li et al.EMNLP 2025 · 2 citations
- Self-Powered LLM Modality Expansion for Large Speech-Text ModelsTengfei Yu, Xuebo Liu, Zhiyi Hou, Liang Ding et al.EMNLP 2024 · 1 citation
- Aya Dataset: An Open-Access Collection for Multilingual Instruction TuningShivalika Singh, Freddie Vargus, Daniel D'souza, Börje Karlsson et al.ACL 2024
