Inverse Constitutional AI: Compressing Preferences into Principles
Arduin Findeis, Timo Kaufmann, Eyke Hüllermeier, Samuel Albanie, Robert Mullins
Abstract
Feedback data is widely used for fine-tuning and evaluating state-of-the-art AI models. Pairwise text preferences, where human or AI annotators select the "better" of two options, are particularly common. Such preferences are used to train (reward) models or to rank models with aggregate statistics. For many applications it is desirable to understand annotator preferences in addition to modelling them -not least because extensive prior work has shown various unintended biases in preference datasets. Yet, preference datasets remain challenging to interpret. Neither black-box reward models nor statistics can answer why one text is preferred over another. Manual interpretation of the numerous (long) response pairs is usually equally infeasible. In this paper, we introduce the Inverse Constitutional AI (ICAI) problem, formulating the interpretation of pairwise text preference data as a compression task. In constitutional AI, a set of principles (a constitution) is used to provide feedback and fine-tune AI models. ICAI inverts this process: given a feedback dataset, we aim to extract a constitution that best enables a large language model (LLM) to reconstruct the original annotations. We propose a corresponding ICAI algorithm and validate its generated constitutions quantitatively based on annotation reconstruction accuracy on several datasets: (a) synthetic feedback data with known principles; (b) AlpacaEval cross-annotated human feedback data; (c) crowdsourced Chatbot Arena data; and (d) PRISM data from diverse demographic groups. As an example application, we further demonstrate the detection of biases in human feedback data. As a short and interpretable representation of the original dataset, generated constitutions have many potential use cases: they may help identify undesirable annotator biases, better understand model performance, scale feedback to unseen data, or assist with adapting AI models to individual user or group preferences. We release the source code for our algorithm and experiments at https://github.com/rdnfn/icai .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- What's In My Human Feedback? Learning Interpretable Descriptions of Preference DataRajiv Movva, Smitha Milli, Sewon Min, Emma PiersonICLR 2026 · 27 citations
- C3AI: Crafting and Evaluating Constitutions for Constitutional AIYara Kyrychenko, Ke Zhou, Edyta Paulina Bogucka, Daniele QuerciaWWW 2025 · 17 citations
- RECAST: Expanding the Boundaries of LLMs' Complex Instruction Following with Multi-Constraint DataZhengkang Guo, Wenhao Liu, Mingchen Xie, Jingwen Xu et al.ICLR 2026 · 13 citations
- Policy Maps: Tools for Guiding the Unbounded Space of LLM BehaviorsMichelle S. Lam, Fred Hohman, Dominik Moritz, Jeffrey P. Bigham et al.UIST 2025 · 4 citations
- First-Person Fairness in ChatbotsTyna Eloundou, Alex Beutel, David G. Robinson, Keren Gu et al.ICLR 2025 · 3 citations
Builds on19
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos et al.ICML 2024 · 1,212 citations
- AlpacaFarm: A Simulation Framework for Methods that Learn from Human FeedbackYann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang et al.NeurIPS 2023 · 948 citations
- LLM Evaluators Recognize and Favor Their Own GenerationsArjun Panickssery, Samuel R. Bowman, Shi FengNeurIPS 2024 · 865 citations
Related papers
- Icon2: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent RegulationQiyuan Chen, Hongsen Huang, Qian Shao, Jiahe Chen et al.EMNLP 2025
- Self-Supervised Alignment with Mutual Information: Learning to Follow Principles without Preference LabelsJan-Philipp Fränken, Eric Zelikman, Rafael Rafailov, Kanishk Gandhi et al.NeurIPS 2024 · 28 citations
- Latent Principle Discovery for Language Model Self-ImprovementKeshav Ramji, Tahira Naseem, Ramón Fernandez AstudilloNeurIPS 2025 · 2 citations
- ELSPR: Evaluator LLM Training Data Self-Purification on Non-Transitive Preferences via Tournament Graph ReconstructionYan Yu, Yilun Liu, Minggui He, Shimin Tao et al.AAAI 2026 · 2 citations
- Mirroring Users: Towards Building Preference-aligned User Simulator with User Feedback in RecommendationTianjun Wei, Huizhong Guo, Yingpeng Du, Zhu Sun et al.ACL 2026 · 4 citations
