Do Slides Help? Multi-modal Context for Automatic Transcription of Conference Talks
Supriti Sinhamahapatra, Jan Niehues
Abstract
State-of-the-art (SOTA) Automatic Speech Recognition (ASR) systems primarily rely on acoustic information while disregarding additional multi-modal context. However, visual information are essential in disambiguation and adaptation. While most work focus on speaker images to handle noise conditions, this work also focuses on integrating presentation slides for the use cases of scientific presentation. In a first step, we create a benchmark for multimodal presentation including an automatic analysis of transcribing domain-specific terminology. Next, we explore methods for augmenting speech models with multi-modal information. We mitigate the lack of datasets with accompanying slides by a suitable approach of data augmentation. Finally, we train a model using the augmented dataset, resulting in a relative reduction in word error rate of approximately 34%, across all words and 35%, for domainspecific terms compared to the baseline model. Our implementation is available 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on5
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- BEATs: Audio Pre-Training with Acoustic TokenizersSanyuan Chen, Yu Wu, Chengyi Wang, Shujie Liu et al.ICML 2023 · 568 citations
- SALMONN: Towards Generic Hearing Abilities for Large Language ModelsChangli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen et al.ICLR 2024 · 557 citations
Related papers
- Speech-Aware Long Context Pruning and Integration for Contextualized Automatic Speech RecognitionYiming Rong, Yixin Zhang, Ziyi Wang, Deyang Jiang et al.AAAI 2026
- Can Visual Context Improve Automatic Speech Recognition for an Embodied Agent?Pradip Pramanick, Chayan SarkarEMNLP 2022 · 5 citations
- MISP-Meeting: A Real-World Dataset with Multimodal Cues for Long-form Meeting Transcription and SummarizationHang Chen, Chao-Han Huck Yang, Jia-Chen Gu, Sabato Marco Siniscalchi et al.ACL 2025
- CIEASR: Contextual Image-Enhanced Automatic Speech Recognition for Improved Homophone DiscriminationZiyi Wang, Yiming Rong, Deyang Jiang, Haoran Wu et al.ACM MM 2024
- VHASR: A Multimodal Speech Recognition System With Vision HotwordsJiliang Hu, Zuchao Li, Ping Wang, Haojun Ai et al.EMNLP 2024
