SafeLab: An Interactive High-Fidelity Benchmark for Embodied Safety in Scientific Robotics
Fengshuo Bai, Yufeng Li, Ruihai Wu, Peishuo Wang, Yuhan Wang, Bernie Zhu, Yuanfei Wang, Tawei Chou, Gao, Runchuan Zhu, Ying Wen, Yaodong Yang, Yuanpei Chen
Abstract
Task success can mask unsafe execution in scientific robotics. On the benchtop, agents must remain safe throughout a rollout rather than merely reach a final goal, because small pose, force, or tilt errors can cause irreversible spillage or equipment damage. Yet prevailing benchmarks emphasize reversible, high-tolerance manipulation, and imitation-trained policies receive no interactive signal to correct execution drift. We introduce SafeLab, a fluid-aware generative bench-Proceedings of the 43 rd International Conference on Machine Learning, Seoul, South Korea. PMLR 306, 2026. Copyright 2026 by the author(s). mark that turns this trajectory-level requirement into an evaluation protocol through verified task synthesis, teleoperation-free expert demonstrations, and dense safety-aware reinforcement learning (RL) feedback. The benchmark provides 64 tasks across 9 manipulation categories, 63 calibrated laboratory assets, and 6,400 expert trajectories. Evaluating five representative policies on SafeLab shows that task success often coexists with safety violations: Safe Success Rate (SSR) trails Success Rate (SR) by more than 30 percentage points in every evaluation domain (liquid handling, instrument actuation, and glassware rearrangement). Bounded residual RL learns execution-level corrections on frozen base policies, improving simulated SSR by 33.1 to 43.0 percentage points, and 50 open-loop physical replays match simulated safety labels in 86% of cases. SafeLab thus treats laboratory readiness as safe execution rather than goal completion alone, and provides a scalable benchmark for screen-
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0ea1b65d-f65d-422a-b718-ffa71d1d104dBuilds on11
- Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement LearningDenis Yarats, Rob Fergus, Alessandro Lazaric, Lerrel PintoICLR 2022 · 457 citations
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic ManipulationTianxing Chen, Zanxin Chen, Baijun Chen, Zijian Cai et al.ICML 2026 · 394 citations
- Efficient Online Reinforcement Learning with Offline DataPhilip J. Ball, Laura Smith, Ilya Kostrikov, Sergey LevineICML 2023 · 326 citations
- Projection-Based Constrained Policy OptimizationTsung-Yen Yang, Justinian Rosca, Karthik Narasimhan, Peter J. RamadgeICLR 2020 · 306 citations
- PiCor: Multi-Task Deep Reinforcement Learning with Policy CorrectionFengshuo Bai, Hongming Zhang, Tianyang Tao, Zhiheng Wu et al.AAAI 2023 · 31 citations
Related papers
- Building a Foundational Guardrail for General Agentic Systems via Synthetic DataYue Huang, Hang Hua, Yujun Zhou, Pengcheng Jing et al.ICLR 2026 · 29 citations
- PACT: Self-Evolving Physical Safety Alignment for Diffusion Policies in Embodied ManipulationLingxuan Wu, Zijian Zhu, Lizhong Wang, Chengyang Ying et al.ICML 2026
- SafeScientist: Enhancing AI Scientist Safety for Risk-Aware Scientific DiscoveryKunlun Zhu, Jiaxun Zhang, Ziheng Qi, Nuoxing Shang et al.EMNLP 2025
- AGENTSAFE: Benchmarking the Safety of Embodied Agents on Hazardous InstructionsZonghao Ying, Le Wang, Yisong Xiao, Jiakai Wang et al.CVPR 2026 · 42 citations
- DRIFT-BENCH: Diagnosing CoopeRative Breakdowns in LLM Agents under Input Faults via Multi-Turn InteractionHan Bao, Zheyuan Zhang, PENGCHENG JING, Zhengqing Yuan et al.ICML 2026
