Reading Between the Tokens: Improving Preference Predictions through Mechanistic Forecasting
Sarah Ball, Simeon Allmendinger, Frauke Kreuter, Niklas Kühl
摘要
Large language models are increasingly used to predict human preferences in both scientific and business endeavors, yet current approaches rely exclusively on analyzing model outputs without considering the underlying mechanisms. Using election forecasting as a test case, we introduce mechanistic forecasting , a method that demonstrates that probing internal model representations offers a fundamentally different---and sometimes more effective--- approach to preference prediction. Examining over 24 million configurations across 7 models, 6 national elections, multiple persona attributes, and prompt variations, we systematically analyze how demographic and ideological information activates latent party-encoding components within the respective models. We find that leveraging this internal knowledge via mechanistic forecasting, opposed to solely relying on surface-level predictions, can improve prediction accuracy. The effects vary across demographic versus opinion-based attributes, political parties, national contexts, and models. Our findings demonstrate that the latent representational structure of LLMs contains systematic, exploitable information about human preferences, establishing a new paradigm for using language models in social science prediction tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- Whose Opinions Do Language Models Reflect?Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee 等ICML 2023 · 被引用 764 次
- Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject StudiesGati V. Aher, Rosa I. Arriaga, Adam Tauman KalaiICML 2023 · 被引用 651 次
- INSIDE: LLMs' Internal States Retain the Power of Hallucination DetectionChao Chen, Kai Liu, Ze Chen, Yi Gu 等ICLR 2024 · 被引用 281 次
- Evaluating Large Language Models in Generating Synthetic HCI Research Data: a Case StudyPerttu Hämäläinen, Mikke Tavast, Anton KunnariCHI 2023 · 被引用 244 次
- A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and ToxicityAndrew Lee, Xiaoyan Bai, Itamar Pres, Martin Wattenberg 等ICML 2024 · 被引用 177 次
相关 Paper
- When Truth Is Overridden: Uncovering the Internal Origins of Sycophancy in Large Language ModelsKeyu Wang, Jin Li, Shu Yang, Zhuoran Zhang 等AAAI 2026 · 被引用 25 次
- Linear Representations of Political Perspective Emerge in Large Language ModelsJunsol Kim, James Evans, Aaron ScheinICLR 2025
- Examining Alignment of Large Language Models through Representative Heuristics: the case of political stereotypesSullam Jeoung, Yubin Ge, Haohan Wang, Jana DiesnerICLR 2025
- What Do Large Language Models Know About Opinions?Erfan Jahanparast, Zhiqing Hong, Serina ChangICLR 2026
- Algorithmic Fidelity of Large Language Models in Generating Synthetic German Public Opinions: A Case StudyBolei Ma, Berk Yoztyurk, Anna-Carolina Haensch, Xinpeng Wang 等ACL 2025 · 被引用 11 次
