MockConf: A Student Interpretation Dataset: Analysis, Word- and Span-level Alignment and Baselines
Dávid Javorský, Ondrej Bojar, François Yvon
Abstract
In simultaneous interpreting, an interpreter renders a source speech into another language with a very short lag, much sooner than sentences are finished. In order to understand and later reproduce this dynamic and complex task automatically, we need dedicated datasets and tools for analysis, monitoring, and evaluation, such as parallel speech corpora, and tools for their automatic annotation. Existing parallel corpora of translated texts and associated alignment algorithms hardly fill this gap, as they fail to model long-range interactions between speech segments or specific types of divergences (e.g., shortening, simplification, functional generalization) between the original and interpreted speeches. In this work, we introduce Mock-Conf, a student interpreting dataset that was collected from Mock Conferences run as part of the students' curriculum. This dataset contains 7 hours of recordings in 5 European languages, transcribed and aligned at the level of spans and words. We further implement and release In-terAlign, a modern web-based annotation tool for parallel word and span annotations on long inputs, suitable for aligning simultaneous interpreting. We propose metrics for the evaluation and a baseline for automatic alignment. Dataset and tools are released to the community.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3cef526e-5a81-41bf-90b5-d0ee409d799bCited by top-tier papers1
Ask how each one uses itBuilds on5
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Detecting and Mitigating Hallucinations in Machine Translation: Model Internal Workings Alone Do Well, Sentence Similarity Even BetterDavid Dale, Elena Voita, Loïc Barrault, Marta R. Costa-jussàACL 2023 · 25 citations
- Optimal Transport for Unsupervised Hallucination Detection in Neural Machine TranslationNuno Miguel Guerreiro, Pierre Colombo, Pablo Piantanida, André F. T. MartinsACL 2023 · 8 citations
- Language-agnostic BERT Sentence EmbeddingFangxiaoyu Feng, Yinfei Yang, Daniel Cer, Naveen Arivazhagan et al.ACL 2022
- Contrastive Learning for Many-to-many Multilingual Neural Machine TranslationXiao Pan, Mingxuan Wang, Liwei Wu, Lei LiACL 2021
Related papers
- Simul-MuST-C: Simultaneous Multilingual Speech Translation Corpus Using Large Language ModelMana Makinae, Yusuke Sakai, Hidetaka Kamigaito, Taro WatanabeEMNLP 2024 · 2 citations
- DEplain: A German Parallel Corpus with Intralingual Translations into Plain Language for Sentence and Document SimplificationRegina Stodden, Omar Momen, Laura KallmeyerACL 2023 · 3 citations
- Simultaneous Interpretation Corpus Construction by Large Language Models in Distant Language PairYusuke Sakai, Mana Makinae, Hidetaka Kamigaito, Taro WatanabeEMNLP 2024 · 3 citations
- StreamAtt: Direct Streaming Speech-to-Text Translation with Attention-based Audio History SelectionSara Papi, Marco Gaido, Matteo Negri, Luisa BentivogliACL 2024
- Learning Adaptive Segmentation Policy for Simultaneous TranslationRuiqing Zhang, Chuanqiang Zhang, Zhongjun He, Hua Wu et al.EMNLP 2020 · 41 citations
