Word Order Does Matter and Shuffled Language Models Know It
Mostafa Abdou, Vinit Ravishankar, Artur Kulmizev, Anders Søgaard
摘要
Recent studies have shown that language models pretrained and/or fine-tuned on randomly permuted sentences exhibit competitive performance on GLUE, putting into question the importance of word order information. Somewhat counter-intuitively, some of these studies also report that position embeddings appear to be crucial for models’ good performance with shuffled text. We probe these language models for word order information and investigate what position embeddings learned from shuffled text encode, showing that these models retain a notion of word order information. We show this is in part due to a subtlety in how shuffling is implemented in previous work – before rather than after subword segmentation. Surprisingly, we find even Language models trained on text shuffled after subword segmentation retain some semblance of information about word order because of the statistical dependencies between sentence length and unigram probabilities. Finally, we show that beyond GLUE, a variety of language understanding tasks do require word order information, often to an extent that cannot be learned through fine-tuning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Premise Order Matters in Reasoning with Large Language ModelsXinyun Chen, Ryan A. Chi, Xuezhi Wang, Denny ZhouICML 2024 · 被引用 59 次
- Mission: Impossible Language ModelsJulie Kallini, Isabel Papadimitriou, Richard Futrell, Kyle Mahowald 等ACL 2024 · 被引用 15 次
- ContextRef: Evaluating Referenceless Metrics for Image Description GenerationElisa Kreiss, Eric Zelikman, Christopher Potts, Nick HaberICLR 2024 · 被引用 6 次
- Order Doesn't Matter, But Reasoning Does: Training LLMs with Order-Centric AugmentationQianxi He, Qianyu He, Jiaqing Liang, Weikang Zhou 等EMNLP 2025
- Interpreting Positional Information in Perspective of Word OrderXilong Zhang, Ruochen Liu, Jin Liu, Xuefeng LiangACL 2023
它引用的顶会 Paper10
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- On the Relationship between Self-Attention and Convolutional LayersJean-Baptiste Cordonnier, Andreas Loukas, Martin JaggiICLR 2020 · 被引用 629 次
- Masked Language Modeling and the Distributional Hypothesis: Order Word Matters Pre-training for LittleKoustuv Sinha, Robin Jia, Dieuwke Hupkes, Joelle Pineau 等EMNLP 2021 · 被引用 177 次
- On Position Embeddings in BERTBenyou Wang, Lifeng Shang, Christina Lioma, Xin Jiang 等ICLR 2021 · 被引用 129 次
- Explicit Regularisation in Gaussian Noise InjectionsAlexander Camuto, Matthew Willetts, Umut Simsekli, Stephen J. Roberts 等NeurIPS 2020 · 被引用 90 次
相关 Paper
- Studying word order through iterative shufflingNikolay Malkin, Sameera Lanka, Pranav Goel, Nebojsa JojicEMNLP 2021 · 被引用 2 次
- On the Importance of Word Order Information in Cross-lingual Sequence LabelingZihan Liu, Genta Indra Winata, Samuel Cahyawijaya, Andrea Madotto 等AAAI 2021 · 被引用 29 次
- Extending the Context of Pretrained LLMs by Dropping Their Positional EmbeddingYoav Gelberg, Koshi Eguchi, Takuya Akiba, Edoardo CetinICLR 2026 · 被引用 13 次
- Computation Mechanism Behind LLM Position GeneralizationChi Han, Heng JiACL 2025
- Fresh in memory: Training-order recency is linearly encoded in language model activationsDmitrii Krasheninnikov, Richard E. Turner, David KruegerICLR 2026 · 被引用 5 次
