Probing LLMs for Multilingual Discourse Generalization Through a Unified Label Set
Florian Eichin, Yang Janet Liu, Barbara Plank, Michael A. Hedderich
Abstract
Discourse understanding is essential for many NLP tasks, yet most existing work remains constrained by framework-dependent discourse representations. This work investigates whether large language models (LLMs) capture discourse knowledge that generalizes across languages and frameworks. We address this question along two dimensions: (1) developing a unified discourse relation label set to facilitate cross-lingual and cross-framework discourse analysis, and (2) probing LLMs to assess whether they encode generalizable discourse abstractions. Using multilingual discourse relation classification as a testbed, we examine a comprehensive set of 23 LLMs of varying sizes and multilingual capabilities. Our results show that LLMs, especially those with multilingual training corpora, can generalize discourse information across languages and frameworks. Further layer-wise analyses reveal that language generalization at the discourse level is most salient in the intermediate layers. Lastly, our error analysis provides an account of challenging relation classes. * Equal contribution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b8f5d6bd-ac70-4744-a3fa-2a9114201636Cited by top-tier papers1
Ask how each one uses itBuilds on9
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Discourse-Aware Neural Extractive Text SummarizationJiacheng Xu, Zhe Gan, Yu Cheng, Jingjing LiuACL 2020 · 264 citations
- Language Modeling Is CompressionGrégoire Delétang, Anian Ruoss, Paul-Ambroise Duquenne, Elliot Catt et al.ICLR 2024 · 243 citations
- The better your Syntax, the better your Semantics? Probing Pretrained Language Models for the English Comparative CorrelativeLeonie Weissweiler, Valentin Hofmann, Abdullatif Köksal, Hinrich SchützeEMNLP 2022 · 13 citations
- Entity Tracking in Language ModelsNajoung Kim, Sebastian SchusterACL 2023 · 9 citations
Related papers
- Improving Implicit Discourse Relation Recognition with Natural Language Explanations from LLMsHeng Wang, Changxing WuAAAI 2026
- Connective Prediction for Implicit Discourse Relation Recognition via Knowledge DistillationHongyi Wu, Hao Zhou, Man Lan, Yuanbin Wu et al.ACL 2023 · 7 citations
- Beyond Chunking: Discourse-Aware Hierarchical Retrieval for Long Document Question AnsweringHuiyao Chen, Yi Yang, Yinghui Li, Meishan Zhang et al.ACL 2026 · 6 citations
- A Language Model-based Generative Classifier for Sentence-level Discourse ParsingYing Zhang, Hidetaka Kamigaito, Manabu OkumuraEMNLP 2021 · 7 citations
- A Label Dependence-Aware Sequence Generation Model for Multi-Level Implicit Discourse Relation RecognitionChangxing Wu, Liuwen Cao, Yubin Ge, Yang Liu et al.AAAI 2022 · 38 citations
