Leveraging Multilingual Training for Authorship Representation: Enhancing Generalization across Languages and Domains
Junghwan Kim, Haotian Zhang, David Jurgens
Abstract
Authorship representation (AR) learning, which models an author's unique writing style, has demonstrated strong performance in authorship attribution tasks. However, prior research has primarily focused on monolingual settings-mostly in English-leaving the potential benefits of multilingual AR models underexplored. We introduce a novel method for multilingual AR learning that incorporates two key innovations: probabilistic content masking, which encourages the model to focus on stylistically indicative words rather than content-specific words, and language-aware batching, which improves contrastive learning by reducing cross-lingual interference. Our model is trained on over 4.5 million authors across 36 languages and 13 domains. It consistently outperforms monolingual baselines in 21 out of 22 non-English languages, achieving an average Recall@8 improvement of 4.85%, with a maximum gain of 15.91% in a single language. Furthermore, it exhibits stronger cross-lingual and cross-domain generalization compared to a monolingual model trained solely on English. Our analysis confirms the effectiveness of both proposed techniques, highlighting their critical roles in the model's improved performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c9f1713e-b9a1-443b-8395-7f53f786ce2bCited by top-tier papers1
Ask how each one uses itBuilds on12
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Few-Shot Detection of Machine-Generated Text using Style RepresentationsRafael A. Rivera Soto, Kailin Koch, Aleem Khan, Barry Y. Chen et al.ICLR 2024 · 49 citations
- Which Features are Learnt by Contrastive Learning? On the Role of Simplicity Bias in Class Collapse and Feature SuppressionYihao Xue, Siddharth Joshi, Eric Gan, Pin-Yu Chen et al.ICML 2023 · 36 citations
Related papers
- Layered Insights: Generalizable Analysis of Human Authorial Style by Leveraging All Transformer LayersMilad Alshomary, Nikhil Reddy Varimalla, Vishal Anand, Smaranda Muresan et al.EMNLP 2025
- Authorship Attribution in Multilingual Machine-Generated TextsLucio La Cava, Dominik Macko, Róbert Móro, Ivan Srba et al.ACL 2026 · 7 citations
- Idiosyncratic but not Arbitrary: Learning Idiolects in Online Registers Reveals Distinctive yet Consistent Individual StylesJian Zhu, David JurgensEMNLP 2021 · 16 citations
- Target-Agnostic Gender-Aware Contrastive Learning for Mitigating Bias in Multilingual Machine TranslationMinwoo Lee, Hyukhun Koh, Kang-il Lee, Dongdong Zhang et al.EMNLP 2023 · 2 citations
- StoryTrans: Non-Parallel Story Author-Style Transfer with Discourse Representations and Content EnhancingXuekai Zhu, Jian Guan, Minlie Huang, Juan LiuACL 2023 · 5 citations
