Idiosyncratic but not Arbitrary: Learning Idiolects in Online Registers Reveals Distinctive yet Consistent Individual Styles
Jian Zhu, David Jurgens
摘要
An individual's variation in writing style is often a function of both social and personal attributes. While structured social variation has been extensively studied, e.g., gender based variation, far less is known about how to characterize individual styles due to their idiosyncratic nature. We introduce a new approach to studying idiolects through a massive crossauthor comparison to identify and encode stylistic features. The neural model achieves strong performance at authorship identification on short texts and through an analogybased probing task, showing that the learned representations exhibit surprising regularities that encode qualitative and quantitative shifts of idiolectal styles. Through text perturbation, we quantify the relative contributions of different linguistic elements to idiolectal variation. Furthermore, we provide a description of idiolects through measuring inter-and intraauthor variation, showing that variation in idiolects is often distinctive yet consistent.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- VendorLink: An NLP approach for Identifying & Linking Vendor Migrants & Potential Aliases on Darknet MarketsVageesh Saxena, Nils Rethmeier, Gijs van Dijck, Gerasimos SpanakisACL 2023 · 被引用 6 次
- Leveraging Multilingual Training for Authorship Representation: Enhancing Generalization across Languages and DomainsJunghwan Kim, Haotian Zhang, David JurgensEMNLP 2025 · 被引用 3 次
- IDTraffickers: An Authorship Attribution Dataset to link and connect Potential Human-Trafficking Operations on Text Escort AdvertisementsVageesh Saxena, Benjamin Bashpole, Gijs van Dijck, Gerasimos SpanakisEMNLP 2023 · 被引用 2 次
- LaMP: When Large Language Models Meet PersonalizationAlireza Salemi, Sheshera Mysore, Michael Bendersky, Hamed ZamaniACL 2024
它引用的顶会 Paper1
相关 Paper
- An Empirical Analysis of the Writing Styles of Persona-Assigned LLMsManuj Malik, Jing Jiang, Kian Ming A. ChaiEMNLP 2024 · 被引用 2 次
- Interacting with Literary Style through Computational ToolsSarah Sterman, Evey Huang, Vivian Liu, Eric PaulosCHI 2020 · 被引用 14 次
- Adapting Language Models for Non-Parallel Author-Stylized RewritingBakhtiyar Syed, Gaurav Verma, Balaji Vasan Srinivasan, Anandhavelu Natarajan 等AAAI 2020 · 被引用 53 次
- Linguistic and Embedding-Based Profiling of Texts Generated by Humans and Large Language ModelsSergio E. Zanotto, Segun AroyehunEMNLP 2025 · 被引用 3 次
- Meta-Tuning LLMs to Leverage Lexical Knowledge for Generalizable Language Style UnderstandingRuohao Guo, Wei Xu, Alan RitterACL 2024 · 被引用 2 次
