Multi-Level Cross-Modal Alignment for Speech Relation Extraction
Liang Zhang, Zhen Yang, Biao Fu, Ziyao Lu, Liangying Shao, Shiyu Liu, Fandong Meng, Jie Zhou, Xiaoli Wang, Jinsong Su
摘要
Speech Relation Extraction (SpeechRE) aims to extract relation triplets from speech data. However, existing studies usually use synthetic speech to train and evaluate SpeechRE models, hindering the further development of SpeechRE due to the disparity between synthetic and real speech. Meanwhile, the modality gap issue, unexplored in SpeechRE, limits the performance of existing models. In this paper, we construct two real SpeechRE datasets to facilitate subsequent researches and propose a Multi-level Cross-modal Alignment Model (MCAM) for SpeechRE. Our model consists of three components: 1) a speech encoder, extracting speech features from the input speech; 2) an alignment adapter, mapping these speech features into a suitable semantic space for the text decoder; and 3) a text decoder, autoregressively generating relation triplets based on the speech features. During training, we first additionally introduce a text encoder to serve as a semantic bridge between the speech encoder and the text decoder, and then train the alignment adapter to align the output features of speech and text encoders at multiple levels. In this way, we can effectively train the alignment adapter to bridge the modality gap between the speech encoder and the text decoder. Experimental results and in-depth analysis on our datasets strongly demonstrate the efficacy of our method. Our source code is available at https://github. com/DeepLearnXMU/SpeechRE-MCAM .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- LLM-OREF: An Open Relation Extraction Framework Based on Large Language ModelsHongyao Tu, Liang Zhang, Yujie Lin, Xin Lin 等EMNLP 2025 · 被引用 2 次
- A Self-Denoising Model for Robust Few-Shot Relation ExtractionLiang Zhang, Yang Zhang, Ziyao Lu, Fandong Meng 等ACL 2025 · 被引用 2 次
- Listening Like Humans: Semantics-Guided Noise-Robust Multimodal Speech RecognitionYan Fang, Jun Chen, Yian Yao, Shuxin Zhong 等ACL 2026
它引用的顶会 Paper17
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- RpBERT: A Text-image Relation Propagation-based BERT Model for Multimodal NERLin Sun, Jiquan Wang, Kai Zhang, Yindu Su 等AAAI 2021 · 被引用 189 次
- Re-TACRED: Addressing Shortcomings of the TACRED DatasetGeorge Stoica, Emmanouil Antonios Platanios, Barnabás PóczosAAAI 2021 · 被引用 146 次
相关 Paper
- Towards relation extraction from speechTongtong Wu, Guitao Wang, Jinming Zhao, Zhaoran Liu 等EMNLP 2022 · 被引用 6 次
- Caption-Aware Multimodal Relation Extraction with Mutual Information MaximizationZefan Zhang, Weiqi Zhang, Yanhui Li, Tian BaiACM MM 2024 · 被引用 9 次
- Entity-centered Cross-document Relation ExtractionFengqi Wang, Fei Li, Hao Fei, Jingye Li 等EMNLP 2022 · 被引用 49 次
- Understanding the Modality Gap: An Empirical Study on the Speech-Text Alignment Mechanism of Large Speech Language ModelsBajian Xiang, Shuaijiang Zhao, Tingwei Guo, Wei ZouEMNLP 2025 · 被引用 6 次
- Aligning Text/Speech Representations from Multimodal Models with MEG Brain Activity During ListeningPadakanti Srijith, Khushbu Pahwa, Radhika Mamidi, Bapi Raju Surampudi 等EMNLP 2025
