From text to talk: Harnessing conversational corpora for humane and diversity-aware language technology
Mark Dingemanse, Andreas Liesenfeld
Abstract
Informal social interaction is the primordial home of human language. Linguistically diverse conversational corpora are an important and largely untapped resource for computational linguistics and language technology. Through the efforts of a worldwide language documentation movement, such corpora are increasingly becoming available. We show how interactional data from 63 languages (26 families) harbours insights about turn-taking, timing, sequential structure and social action, with implications for language technology, natural language understanding, and the design of conversational interfaces. Harnessing linguistically diverse conversational corpora will provide the empirical foundations for flexible, localizable, humane language technologies of the future.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 973ca790-fd18-42c8-aac1-00c6d0d20a43Cited by top-tier papers5
- Multilingual Large Language Models Are Not (Yet) Code-SwitchersRuochen Zhang, Samuel Cahyawijaya, Jan Christian Blaise Cruz, Genta Indra Winata et al.EMNLP 2023 · 19 citations
- Investigating the Representation of Backchannels and Fillers in Fine-tuned Language ModelsYu Wang, Leyi Lao, Langchu Huang, Gabriel Skantze et al.ACL 2026 · 1 citation
- Time is On My Side: Dynamics of Talk-Time Sharing in Video-chat ConversationsKaixiang Zhang, Justine Zhang, Cristian Danescu-Niculescu-MizilCSCW 2025 · 1 citation
- ViClaim: A Multilingual Multilabel Dataset for Automatic Claim Detection in VideosPatrick Giedemann, Pius von Däniken, Jan Milan Deriu, Álvaro Rodrigo et al.EMNLP 2025 · 1 citation
- Alleviating Linguistic and Interactional Anxiety of Non-Native Speakers in Multilingual CommunicationPeinuan Qin, Justin Peng, Zhengtao Xu, Jiting Cheng et al.CSCW 2026
Builds on4
- Climbing towards NLU: On Meaning, Form, and Understanding in the Age of DataEmily M. Bender, Alexander KollerACL 2020 · 914 citations
- The State and Fate of Linguistic Diversity and Inclusion in the NLP WorldPratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali et al.ACL 2020 · 40 citations
- Systematic Inequalities in Language Technology Performance across the World's LanguagesDamián E. Blasi, Antonios Anastasopoulos, Graham NeubigACL 2022
- Changing the World by Changing the DataAnna RogersACL 2021
Related papers
- MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse LanguagesJack FitzGerald, Christopher Hench, Charith Peris, Scott Mackie et al.ACL 2023 · 88 citations
- Fora: A corpus and framework for the study of facilitated dialogueHope Schroeder, Deb Roy, Jad KabbaraACL 2024
- Inter-X: Towards Versatile Human-Human Interaction AnalysisLiang Xu, Xintao Lv, Yichao Yan, Xin Jin et al.CVPR 2024 · 18 citations
- Bloom Library: Multimodal Datasets in 300+ Languages for a Variety of Downstream TasksColin Leong, Joshua Nemecek, Jacob Mansdorfer, Anna Filighera et al.EMNLP 2022 · 3 citations
- Must NLP be Extractive?Steven BirdACL 2024 · 4 citations
