Format as a Prior: Quantifying and Analyzing Bias in LLMs for Heterogeneous Data
Jiacheng Liu, Mayi Xu, Qiankun Pi, Wenli Li, Ming Zhong, Yuanyuan Zhu, Mengchi Liu, Tieyun Qian
Abstract
Large Language Models (LLMs) are increasingly employed in applications that require processing information from heterogeneous formats, including texts, tables, infoboxes, and knowledge graphs. However, systematic biases toward particular formats may undermine LLMs' ability to integrate heterogeneous data impartially, potentially resulting in reasoning errors and increased risks in downstream tasks. Yet it remains unclear whether such biases are systematic, which data-level factors drive them, and what internal mechanisms underlie their emergence.
In this paper, we present the first comprehensive study of format bias in LLMs through a three-stage empirical analysis. The first stage explores the presence and direction of bias across a diverse range of LLMs. The second stage examines how key data-level factors influence these biases. The third stage analyzes how format bias emerges within LLMs' attention patterns and evaluates a lightweight intervention to test its effectiveness. Our results show that format bias is consistent across model families, driven by information richness, structure quality, and representation type, and is closely associated with attention imbalance within the LLMs. Based on these investigations, we identify three future research directions to reduce format bias: enhancing data pre-processing through format repair and normalization, introducing inference-time interventions such as attention re-weighting, and developing format-balanced training corpora. These directions will support the design of more robust and fair heterogeneous data processing systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c9da9e0c-1574-4c3b-949d-84116c548e0fCited by top-tier papers2
- Whose Facts Win? LLM Source Preferences under Knowledge ConflictsJakob Schuster, Vagrant Gautam, Katja MarkertACL 2026 · 3 citations
- Beyond Benchmarks: Toward Causally Faithful Evaluation of Large Language ModelsZhengshuyuan Tian, Wanling Gao, Chuanxin Lan, Chenxi Wang et al.ICML 2026
Builds on6
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Adaptive Chameleon or Stubborn Sloth: Revealing the Behavior of Large Language Models in Knowledge ConflictsJian Xie, Kai Zhang, Jiangjie Chen, Renze Lou et al.ICLR 2024 · 294 citations
- IRCAN: Mitigating Knowledge Conflicts in LLM Generation via Identifying and Reweighting Context-Aware NeuronsDan Shi, Renren Jin, Tianhao Shen, Weilong Dong et al.NeurIPS 2024 · 44 citations
- Knowledge Conflicts for LLMs: A SurveyRongwu Xu, Zehan Qi, Zhijiang Guo, Cunxiang Wang et al.EMNLP 2024 · 38 citations
- Blinded by Generated Contexts: How Language Models Merge Generated and Retrieved Contexts When Knowledge Conflicts?Hexiang Tan, Fei Sun, Wanli Yang, Yuanzhuo Wang et al.ACL 2024
Related papers
- UniBias: Unveiling and Mitigating LLM Bias through Internal Attention and FFN ManipulationHanzhang Zhou, Zijian Feng, Zixiao Zhu, Junlang Qian et al.NeurIPS 2024 · 43 citations
- Attention Speaks Volumes: Localizing and Mitigating Bias in Language ModelsRishabh Adiga, Besmira Nushi, Varun ChandrasekaranACL 2025
- Microstructures and Accuracy of Graph Recall by Large Language ModelsYanbang Wang, Hejie Cui, Jon M. KleinbergNeurIPS 2024 · 3 citations
- Delving into the Reversal Curse: How Far Can Large Language Models Generalize?Zhengkai Lin, Zhihang Fu, Kai Liu, Liang Xie et al.NeurIPS 2024 · 12 citations
- Knowing Bias, Doing Better: Mitigating Social Bias in LLMs via Know-Bias Neuron EnhancementJinhao Pan, Chahat Raj, Anjishnu Mukherjee, Sina Mansouri et al.ICML 2026
