SEED-SET: Scalable Evolving Experimental Design for System-level Ethical Testing
Anjali Parashar, Yingke Li, Eric Yang Yu, Fei Chen, James Neidhoefer, Devesh Upadhyay, Chuchu Fan
Abstract
As autonomous systems such as drones, become increasingly deployed in high-stakes, human-centric domains, it is critical to evaluate the ethical alignment since failure to do so imposes imminent danger to human lives, and long term bias in decision-making. Automated ethical benchmarking of these systems is understudied due to the lack of ubiquitous, well-defined metrics for evaluation, and stakeholder-specific subjectivity, which cannot be modeled analytically. To address these challenges, we propose SEED-SET, a Bayesian experimental design framework that incorporates domain-specific objective evaluations, and subjective value judgments from stakeholders. SEED-SET models both evaluation types separately with hierarchical Gaussian Processes, and uses a novel acquisition strategy to propose interesting test candidates based on learnt qualitative preferences and objectives that align with the stakeholder preferences. We validate our approach for ethical benchmarking of autonomous agents on two applications and find our method to perform the best. Our method provides an interpretable and efficient trade-off between exploration and exploitation, by generating up to optimal test candidates compared to baselines, with improvement in coverage of high dimensional search spaces.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 631f3c1e-bf8d-4600-9419-2a107c8073beBuilds on3
- BoTorch: A Framework for Efficient Monte-Carlo Bayesian OptimizationMaximilian Balandat, Brian Karrer, Daniel R. Jiang, Samuel Daulton et al.NeurIPS 2020 · 686 citations
- Self-Rewarding Language ModelsWeizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Xian Li et al.ICML 2024 · 569 citations
- Eliciting User Preferences for Personalized Multi-Objective Decision Making through Comparative FeedbackHan Shao, Lee Cohen, Avrim Blum, Yishay Mansour et al.NeurIPS 2023 · 10 citations
Related papers
- Model Editing as a Double-Edged Sword: Steering Agent Behavior Toward Beneficence or HarmBaixiang Huang, Zhen Tan, Haoran Wang, Zijie Liu et al.AAAI 2026
- EigenBench: A Comparative Behavioral Measure of Value AlignmentJonathn Chang, Leonhard Piff, Suvadip Sana, Jasmine X. Li et al.ICLR 2026 · 3 citations
- ALI-Agent: Assessing LLMs' Alignment with Human Values via Agent-based EvaluationJingnan Zheng, Han Wang, An Zhang, Tai D. Nguyen et al.NeurIPS 2024 · 61 citations
- Value Alignment VerificationDaniel S. Brown, Jordan Schneider, Anca D. Dragan, Scott NiekumICML 2021 · 41 citations
- Rethinking Objectivity in Clinical AI: A Qualitative Study of Concussion EvaluatorsAli Zaidi, Jessica Jia-Wen Saw, Leigh Fu, Katherine Arneson et al.CSCW 2026
