SafeLab: An Interactive High-Fidelity Benchmark for Embodied Safety in Scientific Robotics
Fengshuo Bai, Yufeng Li, Ruihai Wu, Peishuo Wang, Yuhan Wang, Bernie Zhu, Yuanfei Wang, Tawei Chou, Gao, Runchuan Zhu, Ying Wen, Yaodong Yang, Yuanpei Chen
摘要
Task success can mask unsafe execution in scientific robotics. On the benchtop, agents must remain safe throughout a rollout rather than merely reach a final goal, because small pose, force, or tilt errors can cause irreversible spillage or equipment damage. Yet prevailing benchmarks emphasize reversible, high-tolerance manipulation, and imitation-trained policies receive no interactive signal to correct execution drift. We introduce SafeLab, a fluid-aware generative bench-Proceedings of the 43 rd International Conference on Machine Learning, Seoul, South Korea. PMLR 306, 2026. Copyright 2026 by the author(s). mark that turns this trajectory-level requirement into an evaluation protocol through verified task synthesis, teleoperation-free expert demonstrations, and dense safety-aware reinforcement learning (RL) feedback. The benchmark provides 64 tasks across 9 manipulation categories, 63 calibrated laboratory assets, and 6,400 expert trajectories. Evaluating five representative policies on SafeLab shows that task success often coexists with safety violations: Safe Success Rate (SSR) trails Success Rate (SR) by more than 30 percentage points in every evaluation domain (liquid handling, instrument actuation, and glassware rearrangement). Bounded residual RL learns execution-level corrections on frozen base policies, improving simulated SSR by 33.1 to 43.0 percentage points, and 50 open-loop physical replays match simulated safety labels in 86% of cases. SafeLab thus treats laboratory readiness as safe execution rather than goal completion alone, and provides a scalable benchmark for screen-
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement LearningDenis Yarats, Rob Fergus, Alessandro Lazaric, Lerrel PintoICLR 2022 · 被引用 457 次
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic ManipulationTianxing Chen, Zanxin Chen, Baijun Chen, Zijian Cai 等ICML 2026 · 被引用 394 次
- Efficient Online Reinforcement Learning with Offline DataPhilip J. Ball, Laura Smith, Ilya Kostrikov, Sergey LevineICML 2023 · 被引用 326 次
- Projection-Based Constrained Policy OptimizationTsung-Yen Yang, Justinian Rosca, Karthik Narasimhan, Peter J. RamadgeICLR 2020 · 被引用 306 次
- PiCor: Multi-Task Deep Reinforcement Learning with Policy CorrectionFengshuo Bai, Hongming Zhang, Tianyang Tao, Zhiheng Wu 等AAAI 2023 · 被引用 31 次
相关 Paper
- Building a Foundational Guardrail for General Agentic Systems via Synthetic DataYue Huang, Hang Hua, Yujun Zhou, Pengcheng Jing 等ICLR 2026 · 被引用 29 次
- PACT: Self-Evolving Physical Safety Alignment for Diffusion Policies in Embodied ManipulationLingxuan Wu, Zijian Zhu, Lizhong Wang, Chengyang Ying 等ICML 2026
- SafeScientist: Enhancing AI Scientist Safety for Risk-Aware Scientific DiscoveryKunlun Zhu, Jiaxun Zhang, Ziheng Qi, Nuoxing Shang 等EMNLP 2025
- AGENTSAFE: Benchmarking the Safety of Embodied Agents on Hazardous InstructionsZonghao Ying, Le Wang, Yisong Xiao, Jiakai Wang 等CVPR 2026 · 被引用 42 次
- DRIFT-BENCH: Diagnosing CoopeRative Breakdowns in LLM Agents under Input Faults via Multi-Turn InteractionHan Bao, Zheyuan Zhang, PENGCHENG JING, Zhengqing Yuan 等ICML 2026
