Set-LLM: A Permutation-Invariant LLM
Beni Egressy, Jan Stühmer
摘要
While large language models (LLMs) demonstrate impressive capabilities across numerous applications, their robustness remains a critical concern. This paper is motivated by a specific vulnerability: the order sensitivity of LLMs. This vulnerability manifests itself as the order bias observed when LLMs decide between possible options (for example, a preference for the first option) and the tendency of LLMs to provide different answers when options are reordered. The use cases for this scenario extend beyond the classical case of multiple-choice question answering to the use of LLMs as automated evaluators in AI pipelines, comparing output generated by different models. We introduce Set-LLM, a novel architectural adaptation for pretrained LLMs that enables the processing of mixed set-text inputs with permutation invariance guarantees. The adaptations involve a new attention mask and new positional encodings specifically designed for sets. We provide a theoretical proof of invariance and demonstrate through experiments that Set-LLM can be trained effectively, achieving comparable or improved performance and maintaining the runtime of the original model, while eliminating order sensitivity.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Stable-RAG: Mitigating Retrieval-Permutation-Induced Hallucinations in Retrieval-Augmented GenerationQianchi Zhang, Hainan Zhang, Liang Pang, Hongwei Zheng 等ACL 2026 · 被引用 3 次
- ABCD: All Biases Come DisguisedMateusz Nowak, Xavier Cadet, Peter ChinICML 2026 · 被引用 2 次
- EquiTabPFN: A Target-Permutation Equivariant Prior Fitted NetworkMichael Arbel, David Salinas, Frank HutterNeurIPS 2025 · 被引用 1 次
它引用的顶会 Paper17
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- AlpacaFarm: A Simulation Framework for Methods that Learn from Human FeedbackYann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang 等NeurIPS 2023 · 被引用 948 次
- The Impact of Positional Encoding on Length Generalization in TransformersAmirhossein Kazemnejad, Inkit Padhi, Karthikeyan Natesan Ramamurthy, Payel Das 等NeurIPS 2023 · 被引用 444 次
相关 Paper
- Order-Independence Without Fine TuningReid McIlroy-Young, Katrina Brown, Conlan Olson, Linjun Zhang 等NeurIPS 2024 · 被引用 10 次
- Positional Overload: Positional Debiasing and Context Window Extension for Large Language Models using Set EncodingLukas Kinder, Lukas Edman, Alexander Fraser, Tobias KäferACL 2025
- Fool Your (Vision and) Language Model with Embarrassingly Simple PermutationsYongshuo Zong, Tingyang Yu, Ruchika Chavhan, Bingchen Zhao 等ICML 2024 · 被引用 28 次
- Large Language Models Are Not Robust Multiple Choice SelectorsChujie Zheng, Hao Zhou, Fandong Meng, Jie Zhou 等ICLR 2024 · 被引用 424 次
- RoToR: Towards More Reliable Responses for Order-Invariant InputsSoyoung Yoon, Dongha Ahn, Youngwon Lee, Minkyu Jung 等ACL 2025 · 被引用 1 次
