C3AI: Crafting and Evaluating Constitutions for Constitutional AI
Yara Kyrychenko, Ke Zhou, Edyta Paulina Bogucka, Daniele Quercia
摘要
Constitutional AI (CAI) guides LLM behavior using constitutions, but identifying which principles are most effective for model alignment remains an open challenge. We introduce the C3AI framework (Crafting Constitutions for CAI models), which serves two key functions: (1) selecting and structuring principles to form effective constitutions before fine-tuning; and (2) evaluating whether finetuned CAI models follow these principles in practice. By analyzing principles from AI and psychology, we found that positively framed, behavior-based principles align more closely with human preferences than negatively framed or trait-based principles. In a safety alignment use case, we applied a graph-based principle selection method to refine an existing CAI constitution, improving safety measures while maintaining strong general reasoning capabilities. Interestingly, fine-tuned CAI models performed well on negatively framed principles but struggled with positively framed ones, in contrast to our human alignment results. This highlights a potential gap between principle design and model adherence. Overall, C3AI provides a structured and scalable approach to both crafting and evaluating CAI constitutions. CCS Concepts • Human-centered computing → Collaborative and social computing design and evaluation methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- RECAST: Expanding the Boundaries of LLMs' Complex Instruction Following with Multi-Constraint DataZhengkang Guo, Wenhao Liu, Mingchen Xie, Jingwen Xu 等ICLR 2026 · 被引用 13 次
- The Hall of AI Fears and Hopes: Comparing the Views of AI Influencers and those of Members of the U.S. Public Through an Interactive PlatformGustavo Moreira, Edyta Paulina Bogucka, Marios Constantinides, Daniele QuerciaCHI 2025 · 被引用 5 次
- The State of Multilingual LLM Safety Research: From Measuring The Language Gap To Mitigating ItZheng Xin Yong, Beyza Ermis, Marzieh Fadaee, Stephen H. Bach 等EMNLP 2025 · 被引用 2 次
- There is No War in Ba Sing Se: A Global Analysis of Content Moderation in Large Language ModelsFriedemann Lipphardt, Moonis Ali, Martin Banzer, Anja Feldmann 等NDSS 2026 · 被引用 1 次
它引用的顶会 Paper10
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Whose Opinions Do Language Models Reflect?Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee 等ICML 2023 · 被引用 764 次
- LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation DatasetLianmin Zheng, Wei-Lin Chiang, Ying Sheng, Tianle Li 等ICLR 2024 · 被引用 419 次
- ORPO: Monolithic Preference Optimization without Reference ModelJiwoo Hong, Noah Lee, James ThorneEMNLP 2024 · 被引用 71 次
相关 Paper
- Governance in Motion: Co-evolution of Constitutions and AI models for Scalable SafetyChenhao Huang, Ziyu Shen, Yicong Ren, Huiyuan Zheng 等EMNLP 2025
- Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety RequirementsJingyu Zhang, Ahmed Elgohary, Ahmed Magooda, Daniel Khashabi 等ICLR 2025
- Generative Psycho-Lexical Approach for Constructing Value Systems in Large Language ModelsHaoran Ye, Tianze Zhang, Yuhang Xie, Liyuan Zhang 等ACL 2025 · 被引用 3 次
- Unintended Harms of Value-Aligned LLMs: Psychological and Empirical InsightsSooyung Choi, Jaehyeok Lee, Xiaoyuan Yi, Jing Yao 等ACL 2025
- Towards Better Value Principles for Large Language Model Alignment: A Systematic Evaluation and EnhancementBingbing Xu, Jing Yao, Xiaoyuan Yi, Aishan Maoliniyazi 等ACL 2025 · 被引用 3 次
