C3AI: Crafting and Evaluating Constitutions for Constitutional AI
Yara Kyrychenko, Ke Zhou, Edyta Paulina Bogucka, Daniele Quercia
Abstract
Constitutional AI (CAI) guides LLM behavior using constitutions, but identifying which principles are most effective for model alignment remains an open challenge. We introduce the C3AI framework (Crafting Constitutions for CAI models), which serves two key functions: (1) selecting and structuring principles to form effective constitutions before fine-tuning; and (2) evaluating whether finetuned CAI models follow these principles in practice. By analyzing principles from AI and psychology, we found that positively framed, behavior-based principles align more closely with human preferences than negatively framed or trait-based principles. In a safety alignment use case, we applied a graph-based principle selection method to refine an existing CAI constitution, improving safety measures while maintaining strong general reasoning capabilities. Interestingly, fine-tuned CAI models performed well on negatively framed principles but struggled with positively framed ones, in contrast to our human alignment results. This highlights a potential gap between principle design and model adherence. Overall, C3AI provides a structured and scalable approach to both crafting and evaluating CAI constitutions. CCS Concepts • Human-centered computing → Collaborative and social computing design and evaluation methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 08c73372-cef9-4eec-89c7-2490cd568a52Cited by top-tier papers4
- RECAST: Expanding the Boundaries of LLMs' Complex Instruction Following with Multi-Constraint DataZhengkang Guo, Wenhao Liu, Mingchen Xie, Jingwen Xu et al.ICLR 2026 · 13 citations
- The Hall of AI Fears and Hopes: Comparing the Views of AI Influencers and those of Members of the U.S. Public Through an Interactive PlatformGustavo Moreira, Edyta Paulina Bogucka, Marios Constantinides, Daniele QuerciaCHI 2025 · 5 citations
- The State of Multilingual LLM Safety Research: From Measuring The Language Gap To Mitigating ItZheng Xin Yong, Beyza Ermis, Marzieh Fadaee, Stephen H. Bach et al.EMNLP 2025 · 2 citations
- There is No War in Ba Sing Se: A Global Analysis of Content Moderation in Large Language ModelsFriedemann Lipphardt, Moonis Ali, Martin Banzer, Anja Feldmann et al.NDSS 2026 · 1 citation
Builds on10
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Whose Opinions Do Language Models Reflect?Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee et al.ICML 2023 · 764 citations
- LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation DatasetLianmin Zheng, Wei-Lin Chiang, Ying Sheng, Tianle Li et al.ICLR 2024 · 419 citations
- ORPO: Monolithic Preference Optimization without Reference ModelJiwoo Hong, Noah Lee, James ThorneEMNLP 2024 · 71 citations
Related papers
- Governance in Motion: Co-evolution of Constitutions and AI models for Scalable SafetyChenhao Huang, Ziyu Shen, Yicong Ren, Huiyuan Zheng et al.EMNLP 2025
- Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety RequirementsJingyu Zhang, Ahmed Elgohary, Ahmed Magooda, Daniel Khashabi et al.ICLR 2025
- Generative Psycho-Lexical Approach for Constructing Value Systems in Large Language ModelsHaoran Ye, Tianze Zhang, Yuhang Xie, Liyuan Zhang et al.ACL 2025 · 3 citations
- Unintended Harms of Value-Aligned LLMs: Psychological and Empirical InsightsSooyung Choi, Jaehyeok Lee, Xiaoyuan Yi, Jing Yao et al.ACL 2025
- Towards Better Value Principles for Large Language Model Alignment: A Systematic Evaluation and EnhancementBingbing Xu, Jing Yao, Xiaoyuan Yi, Aishan Maoliniyazi et al.ACL 2025 · 3 citations
