Can Prompt Probe Pretrained Language Models? Understanding the Invisible Risks from a Causal View
Boxi Cao, Hongyu Lin, Xianpei Han, Fangchao Liu, Le Sun
Abstract
Prompt-based probing has been widely used in evaluating the abilities of pretrained language models (PLMs). Unfortunately, recent studies have discovered such an evaluation may be inaccurate, inconsistent and unreliable. Furthermore, the lack of understanding its inner workings, combined with its wide applicability, has the potential to lead to unforeseen risks for evaluating and applying PLMs in real-world applications. To discover, understand and quantify the risks, this paper investigates the prompt-based probing from a causal view, highlights three critical biases which could induce biased results and conclusions, and proposes to conduct debiasing via causal intervention. This paper provides valuable insights for the design of unbiased datasets, better probing frameworks and more reliable evaluations of pretrained language models. Furthermore, our conclusions also echo that we need to rethink the criteria for identifying better pretrained language models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 602ca8ed-b6f9-417e-8149-85bfbb9da5f8Cited by top-tier papers8
- Causality-aware Concept Extraction based on Knowledge-guided PromptingSiyu Yuan, Deqing Yang, Jinxi Liu, Shuyu Tian et al.ACL 2023 · 7 citations
- A Causal Explainable Guardrails for Large Language ModelsZhixuan Chu, Yan Wang, Longfei Li, Zhibo Wang et al.CCS 2024 · 5 citations
- Boosting Resilience of Large Language Models through Causality-Driven Robust OptimizationXiaoling Zhou, Mingjie Zhang, Zhemg Lee, Yuncheng Hua et al.NeurIPS 2025 · 5 citations
- Neuro-Symbolic Procedural Planning with Commonsense PromptingYujie Lu, Weixi Feng, Wanrong Zhu, Wenda Xu et al.ICLR 2023 · 3 citations
- Does the Correctness of Factual Knowledge Matter for Factual Knowledge-Enhanced Pre-trained Language Models?Boxi Cao, Qiaoyu Tang, Hongyu Lin, Xianpei Han et al.EMNLP 2023
Builds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated PromptsTaylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace et al.EMNLP 2020 · 1,162 citations
- Investigating Gender Bias in Language Models Using Causal Mediation AnalysisJesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian et al.NeurIPS 2020 · 851 citations
- Evaluating Commonsense in Pre-Trained Language ModelsXuhui Zhou, Yue Zhang, Leyang Cui, Dandan HuangAAAI 2020 · 198 citations
Related papers
- Unbiased Evaluation of Large Language Models from a Causal PerspectiveMeilin Chen, Jian Tian, Liang Ma, Di Xie et al.ICML 2025
- Causal Prompting: Debiasing Large Language Model Prompting Based on Front-Door AdjustmentCongzhi Zhang, Linhai Zhang, Jialong Wu, Yulan He et al.AAAI 2025 · 42 citations
- Beyond Benchmarks: Toward Causally Faithful Evaluation of Large Language ModelsZhengshuyuan Tian, Wanling Gao, Chuanxin Lan, Chenxi Wang et al.ICML 2026
- Prompting is not a substitute for probability measurements in large language modelsJennifer Hu, Roger LevyEMNLP 2023 · 31 citations
- Causal-Debias: Unifying Debiasing in Pretrained Language Models and Fine-tuning via Causal Invariant LearningFan Zhou, Yuzhou Mao, Liu Yu, Yi Yang et al.ACL 2023 · 21 citations
