Style is NOT a single variable: Case Studies for Cross-Stylistic Language Understanding
Dongyeop Kang, Eduard H. Hovy
Abstract
Every natural text is written in some style. Style is formed by a complex combination of different stylistic factors, including formality markers, emotions, metaphors, etc. One cannot form a complete understanding of a text without considering these factors. The factors combine and co-vary in complex ways to form styles. Studying the nature of the covarying combinations sheds light on stylistic language in general, sometimes called crossstyle language understanding. This paper provides the benchmark corpus (XSLUE) that combines existing datasets and collects a new one for sentence-level cross-style language understanding and evaluation. The benchmark contains text in 15 different styles under the proposed four theoretical groupings: figurative, personal, affective, and interpersonal groups. For valid evaluation, we collect an additional diagnostic set by annotating all 15 styles on the same text. Using XSLUE, we propose three interesting crossstyle applications in classification, correlation, and generation. First, our proposed crossstyle classifier trained with multiple styles together helps improve overall classification performance against individually-trained style classifiers. Second, our study shows that some styles are highly dependent on each other in human-written text. Finally, we find that combinations of some contradictive styles likely generate stylistically less appropriate text. We believe our benchmark and case studies help explore interesting future directions for crossstyle research. The preprocessed datasets and code are publicly available. 1 * * This work was done while DK was at CMU. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Do LLMs Understand Social Knowledge? Evaluating the Sociability of Large Language Models with SocKET BenchmarkMinje Choi, Jiaxin Pei, Sagar Kumar, Chang Shu et al.EMNLP 2023 · 36 citations
- When Style Breaks Safety: Defending LLMs Against Superficial Style AlignmentYuxin Xiao, Sana Tonekaboni, Walter Gerych, Vinith Menon Suriyakumar et al.ICLR 2026 · 8 citations
- Does It Capture STEL? A Modular, Similarity-based Linguistic Style Evaluation FrameworkAnna Wegmann, Dong NguyenEMNLP 2021 · 7 citations
- Dynamic Multi-Reward Weighting for Multi-Style Controllable GenerationKarin de Langis, Ryan Koo, Dongyeop KangEMNLP 2024 · 3 citations
- Comparing Styles across LanguagesShreya Havaldar, Matthew Pressimone, Eric Wong, Lyle H. UngarEMNLP 2023 · 1 citation
Builds on1
Related papers
- XGLUE: A New Benchmark Datasetfor Cross-lingual Pre-training, Understanding and GenerationYaobo Liang, Nan Duan, Yeyun Gong, Ning Wu et al.EMNLP 2020 · 232 citations
- SciNLI: A Corpus for Natural Language Inference on Scientific TextMobashir Sadat, Cornelia CarageaACL 2022 · 41 citations
- Interacting with Literary Style through Computational ToolsSarah Sterman, Evey Huang, Vivian Liu, Eric PaulosCHI 2020 · 14 citations
- MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning BenchmarkDingdong Wang, Junan Li, Jincenzi Wu, Dongchao Yang et al.ICLR 2026 · 143 citations
- Expertise Style Transfer: A New Task Towards Better Communication between Experts and LaymenYixin Cao, Ruihao Shui, Liangming Pan, Min-Yen Kan et al.ACL 2020 · 50 citations
