Enforcing Axioms for AI Alignment under Loss-Based Rules
Alexandros Hollender, Sonja Kraiczy
Abstract
Recent alignment methods for large language models, most notably reinforcement learning from human feedback (RLHF), often train an auxiliary reward model to minimize a loss function on binary preference data over model responses. We study a theoretical setting inspired by principle-guided methods such as Constitutional AI, in which a small set of principles (e.g., helpfulness, toxicity) act as “voters” that guide binary comparisons---such as preferring the less toxic response. We model these principles as linear directions in an embedding space of responses, a simplifying assumption motivated by the Linear Representation Hypothesis---concepts are linear directions in representation-space---a useful first-order approximation in practice. In this linear social choice model, Ge et al. (2024) showed that an optimal linear reward model can violate Pareto optimality (PO): From the principles-as-voters lens, this means a response A can be less helpful and more toxic than B, yet still receive a higher reward. We analyze axiomatic violations in the linear social choice setting and probe the robustness of negative results under realistic assumptions. We show that added expressivity does not resolve the issue: polynomial reward models can still fail PO. We then offer a pragmatic alternative showing that when the data uniformly covers the embedding space, broad classes of loss-based rules in the limit exactly recover the axiomatic guarantees. This yields a recipe for constitutional-style alignment with provable guarantees: enforce balanced coverage via dataset design to restore axiomatic guarantees without abandoning standard training pipelines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 06c5b1d4-a167-4f7d-b26f-c7411304a163Builds on9
- Principle-Driven Self-Alignment of Language Models from Scratch with Minimal Human SupervisionZhiqing Sun, Yikang Shen, Qinhong Zhou, Hongxin Zhang et al.NeurIPS 2023 · 463 citations
- The Linear Representation Hypothesis and the Geometry of Large Language ModelsKiho Park, Yo Joong Choe, Victor VeitchICML 2024 · 461 citations
- ULTRAFEEDBACK: Boosting Language Models with Scaled AI FeedbackGanqu Cui, Lifan Yuan, Ning Ding, Guanming Yao et al.ICML 2024 · 286 citations
- Nash Learning from Human FeedbackRémi Munos, Michal Valko, Daniele Calandriello, Mohammad Gheshlaghi Azar et al.ICML 2024 · 212 citations
- Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHFAnand Siththaranjan, Cassidy Laidlaw, Dylan Hadfield-MenellICLR 2024 · 112 citations
Related papers
- Axiomatic Preference Modeling for Longform Question AnsweringCorby Rosset, Guoqing Zheng, Victor Dibia, Ahmed Awadallah et al.EMNLP 2023 · 2 citations
- MaxMin-RLHF: Alignment with Diverse Human PreferencesSouradip Chakraborty, Jiahao Qiu, Hui Yuan, Alec Koppel et al.ICML 2024 · 104 citations
- Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences?Paul Gölz, Nika Haghtalab, Kunhe YangNeurIPS 2025 · 29 citations
- Clone-Robust AI AlignmentAriel D. Procaccia, Benjamin Schiffer, Shirley ZhangICML 2025
- More RLHF, More Trust? On The Impact of Preference Alignment On TrustworthinessAaron Jiaxun Li, Satyapriya Krishna, Himabindu LakkarajuICLR 2025
