Inverse Constitutional AI: Compressing Preferences into Principles
Arduin Findeis, Timo Kaufmann, Eyke Hüllermeier, Samuel Albanie, Robert Mullins
摘要
Feedback data is widely used for fine-tuning and evaluating state-of-the-art AI models. Pairwise text preferences, where human or AI annotators select the "better" of two options, are particularly common. Such preferences are used to train (reward) models or to rank models with aggregate statistics. For many applications it is desirable to understand annotator preferences in addition to modelling them -not least because extensive prior work has shown various unintended biases in preference datasets. Yet, preference datasets remain challenging to interpret. Neither black-box reward models nor statistics can answer why one text is preferred over another. Manual interpretation of the numerous (long) response pairs is usually equally infeasible. In this paper, we introduce the Inverse Constitutional AI (ICAI) problem, formulating the interpretation of pairwise text preference data as a compression task. In constitutional AI, a set of principles (a constitution) is used to provide feedback and fine-tune AI models. ICAI inverts this process: given a feedback dataset, we aim to extract a constitution that best enables a large language model (LLM) to reconstruct the original annotations. We propose a corresponding ICAI algorithm and validate its generated constitutions quantitatively based on annotation reconstruction accuracy on several datasets: (a) synthetic feedback data with known principles; (b) AlpacaEval cross-annotated human feedback data; (c) crowdsourced Chatbot Arena data; and (d) PRISM data from diverse demographic groups. As an example application, we further demonstrate the detection of biases in human feedback data. As a short and interpretable representation of the original dataset, generated constitutions have many potential use cases: they may help identify undesirable annotator biases, better understand model performance, scale feedback to unseen data, or assist with adapting AI models to individual user or group preferences. We release the source code for our algorithm and experiments at https://github.com/rdnfn/icai .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- What's In My Human Feedback? Learning Interpretable Descriptions of Preference DataRajiv Movva, Smitha Milli, Sewon Min, Emma PiersonICLR 2026 · 被引用 27 次
- C3AI: Crafting and Evaluating Constitutions for Constitutional AIYara Kyrychenko, Ke Zhou, Edyta Paulina Bogucka, Daniele QuerciaWWW 2025 · 被引用 17 次
- RECAST: Expanding the Boundaries of LLMs' Complex Instruction Following with Multi-Constraint DataZhengkang Guo, Wenhao Liu, Mingchen Xie, Jingwen Xu 等ICLR 2026 · 被引用 13 次
- Policy Maps: Tools for Guiding the Unbounded Space of LLM BehaviorsMichelle S. Lam, Fred Hohman, Dominik Moritz, Jeffrey P. Bigham 等UIST 2025 · 被引用 4 次
- First-Person Fairness in ChatbotsTyna Eloundou, Alex Beutel, David G. Robinson, Keren Gu 等ICLR 2025 · 被引用 3 次
它引用的顶会 Paper19
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos 等ICML 2024 · 被引用 1,212 次
- AlpacaFarm: A Simulation Framework for Methods that Learn from Human FeedbackYann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang 等NeurIPS 2023 · 被引用 948 次
- LLM Evaluators Recognize and Favor Their Own GenerationsArjun Panickssery, Samuel R. Bowman, Shi FengNeurIPS 2024 · 被引用 865 次
相关 Paper
- Icon2: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent RegulationQiyuan Chen, Hongsen Huang, Qian Shao, Jiahe Chen 等EMNLP 2025
- Self-Supervised Alignment with Mutual Information: Learning to Follow Principles without Preference LabelsJan-Philipp Fränken, Eric Zelikman, Rafael Rafailov, Kanishk Gandhi 等NeurIPS 2024 · 被引用 28 次
- Latent Principle Discovery for Language Model Self-ImprovementKeshav Ramji, Tahira Naseem, Ramón Fernandez AstudilloNeurIPS 2025 · 被引用 2 次
- ELSPR: Evaluator LLM Training Data Self-Purification on Non-Transitive Preferences via Tournament Graph ReconstructionYan Yu, Yilun Liu, Minggui He, Shimin Tao 等AAAI 2026 · 被引用 2 次
- Mirroring Users: Towards Building Preference-aligned User Simulator with User Feedback in RecommendationTianjun Wei, Huizhong Guo, Yingpeng Du, Zhu Sun 等ACL 2026 · 被引用 4 次
