Do Slides Help? Multi-modal Context for Automatic Transcription of Conference Talks
Supriti Sinhamahapatra, Jan Niehues
摘要
State-of-the-art (SOTA) Automatic Speech Recognition (ASR) systems primarily rely on acoustic information while disregarding additional multi-modal context. However, visual information are essential in disambiguation and adaptation. While most work focus on speaker images to handle noise conditions, this work also focuses on integrating presentation slides for the use cases of scientific presentation. In a first step, we create a benchmark for multimodal presentation including an automatic analysis of transcribing domain-specific terminology. Next, we explore methods for augmenting speech models with multi-modal information. We mitigate the lack of datasets with accompanying slides by a suitable approach of data augmentation. Finally, we train a model using the augmented dataset, resulting in a relative reduction in word error rate of approximately 34%, across all words and 35%, for domainspecific terms compared to the baseline model. Our implementation is available 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper5
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- BEATs: Audio Pre-Training with Acoustic TokenizersSanyuan Chen, Yu Wu, Chengyi Wang, Shujie Liu 等ICML 2023 · 被引用 568 次
- SALMONN: Towards Generic Hearing Abilities for Large Language ModelsChangli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen 等ICLR 2024 · 被引用 557 次
相关 Paper
- Speech-Aware Long Context Pruning and Integration for Contextualized Automatic Speech RecognitionYiming Rong, Yixin Zhang, Ziyi Wang, Deyang Jiang 等AAAI 2026
- Can Visual Context Improve Automatic Speech Recognition for an Embodied Agent?Pradip Pramanick, Chayan SarkarEMNLP 2022 · 被引用 5 次
- MISP-Meeting: A Real-World Dataset with Multimodal Cues for Long-form Meeting Transcription and SummarizationHang Chen, Chao-Han Huck Yang, Jia-Chen Gu, Sabato Marco Siniscalchi 等ACL 2025
- CIEASR: Contextual Image-Enhanced Automatic Speech Recognition for Improved Homophone DiscriminationZiyi Wang, Yiming Rong, Deyang Jiang, Haoran Wu 等ACM MM 2024
- VHASR: A Multimodal Speech Recognition System With Vision HotwordsJiliang Hu, Zuchao Li, Ping Wang, Haojun Ai 等EMNLP 2024
