A Peek into Token Bias: Large Language Models Are Not Yet Genuine Reasoners
Bowen Jiang, Yangxinyu Xie, Zhuoqun Hao, Xiaomeng Wang, Tanwi Mallick, Weijie Su, Camillo J. Taylor, Dan Roth
摘要
This study introduces a hypothesis-testing framework to assess whether large language models (LLMs) possess genuine reasoning abilities or primarily depend on token bias. We go beyond evaluating LLMs on accuracy; rather, we aim to investigate their token bias in solving logical reasoning tasks. Specifically, we develop carefully controlled synthetic datasets, featuring conjunction fallacy and syllogistic problems. Our framework outlines a list of hypotheses where token biases are readily identifiable, with all null hypotheses assuming genuine reasoning capabilities of LLMs. The findings in this study suggest, with statistical guarantee, that most LLMs still struggle with logical reasoning. While they may perform well on classic problems, their success largely depends on recognizing superficial patterns with strong token bias, thereby raising concerns about their actual reasoning and generalization abilities. Codes and data are open-sourced at https://github.com/bowenupenn/llm_token_bias .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- Teaching Language Models to Evolve with Users: Dynamic Profile Modeling for Personalized AlignmentWeixiang Zhao, Xingyu Sui, Yulin Hu, Jiahe Guo 等NeurIPS 2025 · 被引用 34 次
- ASyMOB: Algebraic Symbolic Mathematical Operations BenchmarkMichael Shalyt, Rotem Elimelech, Ido KaminerICML 2026 · 被引用 7 次
- Quadratic Coreset Selection: Certifying and Reconciling Sequence and Token Mining for Efficient Instruction TuningZiliang Chen, Yongsen Zheng, Zhao-Rong Lai, Zhanfu Yang 等NeurIPS 2025 · 被引用 4 次
- LINGOLY-TOO: Disentangling Reasoning from Knowledge with Templatised Orthographic ObfuscationJude Khouja, Lingyi Yang, Karolina Korgul, Simeon Hellsten 等ICLR 2026 · 被引用 4 次
- TRIM: Token-wise Attention-Derived Saliency for Data-Efficient Instruction TuningManish Nagaraj, Sakshi Choudhary, Utkarsh Saxena, Deepak Ravikumar 等ICML 2026 · 被引用 2 次
它引用的顶会 Paper12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 被引用 1,792 次
- Graph of Thoughts: Solving Elaborate Problems with Large Language ModelsMaciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger 等AAAI 2024 · 被引用 1,292 次
- Large Language Models Cannot Self-Correct Reasoning YetJie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng 等ICLR 2024 · 被引用 858 次
相关 Paper
- Logical forms complement probability in understanding language model (and human) performanceYixuan Wang, Freda ShiACL 2025 · 被引用 2 次
- Test of Time: A Benchmark for Evaluating LLMs on Temporal ReasoningBahare Fatemi, Mehran Kazemi, Anton Tsitsulin, Karishma Malkan 等ICLR 2025 · 被引用 2 次
- Conditional and Modal Reasoning in Large Language ModelsWesley H. Holliday, Matthew Mandelkern, Cedegao ZhangEMNLP 2024 · 被引用 5 次
- SATBench: Benchmarking LLMs' Logical Reasoning via Automated Puzzle Generation from SAT FormulasAnjiang Wei, Yuheng Wu, Yingjia Wan, Tarun Suresh 等EMNLP 2025 · 被引用 1 次
- LogicBench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language ModelsMihir Parmar, Nisarg Patel, Neeraj Varshney, Mutsumi Nakamura 等ACL 2024
