Towards Cognitively-Faithful Decision-Making Models to Improve AI Alignment
Cyrus Cousins, Vijay Keswani, Vincent Conitzer, Hoda Heidari, Jana Schaich Borg, Walter Sinnott-Armstrong
Abstract
Recent AI trends seek to align AI models to learned human-centric objectives, such as personal preferences, utility, or societal values. Using standard preference elicitation methods, researchers and practitioners build models of human decisions and judgments, to which AI models are aligned. However, standard elicitation methods often fail to capture the true cognitive processes behind human decision making, such as the use of heuristics or simplifying structured thought patterns. To address this limitation, we take an axiomatic approach to learning cognitively faithful decision processes from pairwise comparisons. Building on the vast literature characterizing cognitive processes that contribute to human decision-making and pairwise comparisons, we derive a class of models in which individual features are first processed with learned rules, then aggregated via a fixed rule, such as the Bradley-Terry rule, to produce a decision. This structured processing of information ensures that such models are realistic and feasible candidates to represent underlying human decision-making processes. We demonstrate the efficacy of this modeling approach by learning interpretable models of human decision making in a kidney allocation task, and show that our proposed models match or surpass the accuracy of prior models of human pairwise decision-making.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 723df97c-5a94-4f3f-b711-1d481dba1bacCited by top-tier papers1
Ask how each one uses itBuilds on9
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- The Effects of Reward Misspecification: Mapping and Mitigating Misaligned ModelsAlexander Pan, Kush Bhatia, Jacob SteinhardtICLR 2022 · 293 citations
- What is Human-Centered about Human-Centered AI? A Map of the Research LandscapeTara Capel, Margot BreretonCHI 2023 · 218 citations
- Consequences of Misaligned AISimon Zhuang, Dylan Hadfield-MenellNeurIPS 2020 · 120 citations
Related papers
- Can AI Model the Complexities of Human Moral Decision-making? A Qualitative Study of Kidney Allocation DecisionsVijay Keswani, Vincent Conitzer, Walter Sinnott-Armstrong, Breanna K. Nguyen et al.CHI 2025 · 11 citations
- Preference Learning of Latent Decision Utilities with a Human-like Model of Preferential ChoiceSebastiaan De Peuter, Shibei Zhu, Yujia Guo, Andrew Howes et al.NeurIPS 2024 · 6 citations
- Axioms for AI Alignment from Human FeedbackLuise Ge, Daniel Halpern, Evi Micha, Ariel D. Procaccia et al.NeurIPS 2024 · 64 citations
- Human Reliance on Machine Learning Models When Performance Feedback is Limited: Heuristics and RisksZhuoran Lu, Ming YinCHI 2021 · 123 citations
- Apparently Irrational Choice as Optimal Sequential Decision MakingHaiyang Chen, Hyung Jin Chang, Andrew HowesAAAI 2021 · 10 citations
