SEED-SET: Scalable Evolving Experimental Design for System-level Ethical Testing
Anjali Parashar, Yingke Li, Eric Yang Yu, Fei Chen, James Neidhoefer, Devesh Upadhyay, Chuchu Fan
摘要
As autonomous systems such as drones, become increasingly deployed in high-stakes, human-centric domains, it is critical to evaluate the ethical alignment since failure to do so imposes imminent danger to human lives, and long term bias in decision-making. Automated ethical benchmarking of these systems is understudied due to the lack of ubiquitous, well-defined metrics for evaluation, and stakeholder-specific subjectivity, which cannot be modeled analytically. To address these challenges, we propose SEED-SET, a Bayesian experimental design framework that incorporates domain-specific objective evaluations, and subjective value judgments from stakeholders. SEED-SET models both evaluation types separately with hierarchical Gaussian Processes, and uses a novel acquisition strategy to propose interesting test candidates based on learnt qualitative preferences and objectives that align with the stakeholder preferences. We validate our approach for ethical benchmarking of autonomous agents on two applications and find our method to perform the best. Our method provides an interpretable and efficient trade-off between exploration and exploitation, by generating up to optimal test candidates compared to baselines, with improvement in coverage of high dimensional search spaces.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- BoTorch: A Framework for Efficient Monte-Carlo Bayesian OptimizationMaximilian Balandat, Brian Karrer, Daniel R. Jiang, Samuel Daulton 等NeurIPS 2020 · 被引用 686 次
- Self-Rewarding Language ModelsWeizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Xian Li 等ICML 2024 · 被引用 569 次
- Eliciting User Preferences for Personalized Multi-Objective Decision Making through Comparative FeedbackHan Shao, Lee Cohen, Avrim Blum, Yishay Mansour 等NeurIPS 2023 · 被引用 10 次
相关 Paper
- Model Editing as a Double-Edged Sword: Steering Agent Behavior Toward Beneficence or HarmBaixiang Huang, Zhen Tan, Haoran Wang, Zijie Liu 等AAAI 2026
- EigenBench: A Comparative Behavioral Measure of Value AlignmentJonathn Chang, Leonhard Piff, Suvadip Sana, Jasmine X. Li 等ICLR 2026 · 被引用 3 次
- ALI-Agent: Assessing LLMs' Alignment with Human Values via Agent-based EvaluationJingnan Zheng, Han Wang, An Zhang, Tai D. Nguyen 等NeurIPS 2024 · 被引用 61 次
- Value Alignment VerificationDaniel S. Brown, Jordan Schneider, Anca D. Dragan, Scott NiekumICML 2021 · 被引用 41 次
- Rethinking Objectivity in Clinical AI: A Qualitative Study of Concussion EvaluatorsAli Zaidi, Jessica Jia-Wen Saw, Leigh Fu, Katherine Arneson 等CSCW 2026
