Word Order Does Matter and Shuffled Language Models Know It
Mostafa Abdou, Vinit Ravishankar, Artur Kulmizev, Anders Søgaard
Abstract
Recent studies have shown that language models pretrained and/or fine-tuned on randomly permuted sentences exhibit competitive performance on GLUE, putting into question the importance of word order information. Somewhat counter-intuitively, some of these studies also report that position embeddings appear to be crucial for models’ good performance with shuffled text. We probe these language models for word order information and investigate what position embeddings learned from shuffled text encode, showing that these models retain a notion of word order information. We show this is in part due to a subtlety in how shuffling is implemented in previous work – before rather than after subword segmentation. Surprisingly, we find even Language models trained on text shuffled after subword segmentation retain some semblance of information about word order because of the statistical dependencies between sentence length and unigram probabilities. Finally, we show that beyond GLUE, a variety of language understanding tasks do require word order information, often to an extent that cannot be learned through fine-tuning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4e84bbac-bc5e-4d27-9b09-616aefad3f78Cited by top-tier papers7
- Premise Order Matters in Reasoning with Large Language ModelsXinyun Chen, Ryan A. Chi, Xuezhi Wang, Denny ZhouICML 2024 · 59 citations
- Mission: Impossible Language ModelsJulie Kallini, Isabel Papadimitriou, Richard Futrell, Kyle Mahowald et al.ACL 2024 · 15 citations
- ContextRef: Evaluating Referenceless Metrics for Image Description GenerationElisa Kreiss, Eric Zelikman, Christopher Potts, Nick HaberICLR 2024 · 6 citations
- Order Doesn't Matter, But Reasoning Does: Training LLMs with Order-Centric AugmentationQianxi He, Qianyu He, Jiaqing Liang, Weikang Zhou et al.EMNLP 2025
- Interpreting Positional Information in Perspective of Word OrderXilong Zhang, Ruochen Liu, Jin Liu, Xuefeng LiangACL 2023
Builds on10
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- On the Relationship between Self-Attention and Convolutional LayersJean-Baptiste Cordonnier, Andreas Loukas, Martin JaggiICLR 2020 · 629 citations
- Masked Language Modeling and the Distributional Hypothesis: Order Word Matters Pre-training for LittleKoustuv Sinha, Robin Jia, Dieuwke Hupkes, Joelle Pineau et al.EMNLP 2021 · 177 citations
- On Position Embeddings in BERTBenyou Wang, Lifeng Shang, Christina Lioma, Xin Jiang et al.ICLR 2021 · 129 citations
- Explicit Regularisation in Gaussian Noise InjectionsAlexander Camuto, Matthew Willetts, Umut Simsekli, Stephen J. Roberts et al.NeurIPS 2020 · 90 citations
Related papers
- Studying word order through iterative shufflingNikolay Malkin, Sameera Lanka, Pranav Goel, Nebojsa JojicEMNLP 2021 · 2 citations
- On the Importance of Word Order Information in Cross-lingual Sequence LabelingZihan Liu, Genta Indra Winata, Samuel Cahyawijaya, Andrea Madotto et al.AAAI 2021 · 29 citations
- Extending the Context of Pretrained LLMs by Dropping Their Positional EmbeddingYoav Gelberg, Koshi Eguchi, Takuya Akiba, Edoardo CetinICLR 2026 · 13 citations
- Computation Mechanism Behind LLM Position GeneralizationChi Han, Heng JiACL 2025
- Fresh in memory: Training-order recency is linearly encoded in language model activationsDmitrii Krasheninnikov, Richard E. Turner, David KruegerICLR 2026 · 5 citations
