ICML 2026 SafeLab - Heungwoo/research GitHub Wiki
SafeLab: An Interactive High-Fidelity Benchmark for Embodied Safety in Scientific Robotics — Zero-tolerance safety for robots in the chemistry lab
Venue: ICML 2026 (Poster) Category: Benchmark (Embodied Safety / Scientific Robotics) Affiliations: Fengshuo Bai, Yufeng Li, Ruihai Wu, Peishuo Wang, Yuhan Wang, Bernie Zhu, Yuanfei Wang, Tawei Chou, Runchuan Zhu, Ying Wen, Yaodong Yang, Yuanpei Chen
Laboratory automation driven by scientific embodied agents is framed as a critical frontier for modern laboratories — but conventional robotic benchmarks make a forgiving assumption that lab settings do not share. Standard manipulation benchmarks allow reversible failures: a dropped block can simply be picked up again. A chemistry lab demands precision with zero tolerance for errors, because a single mistake can trigger chemical hazards or equipment damage. The authors further identify a structural weakness in current Vision-Language-Action (VLA) models: they rely on static imitation learning and therefore cannot recover from execution drift, so small errors compound over the long horizons typical of precision-critical lab tasks.
SafeLab is a generative simulation benchmark designed for the full lifecycle of safe robot learning, grounded in a high-fidelity chemistry lab. It integrates three components:
- An LLM task-synthesis engine that generates the lab manipulation tasks.
- An automated expert that collects demonstration trajectories for those tasks.
- An interactive RL environment that supports continuous reinforcement-learning refinement of the policy after imitation pretraining.
The team released 6,000+ complex trajectories for evaluation. The interactive environment is the key differentiator: instead of only scoring a frozen policy, it lets agents act, observe the consequences of unsafe behavior, and learn active error correction — addressing the execution-drift failure mode that static imitation learning cannot.
flowchart LR
A[LLM task-synthesis engine] --> B[Automated expert<br/>demo collection]
B --> C[6000+ trajectories]
C --> D[Imitation pretraining]
D --> E[Interactive RL environment<br/>continuous refinement]
E --> F[Safer embodied agent<br/>active error correction]
Evaluating current embodied agents on SafeLab, the authors find that they fail significantly under the safety constraints the benchmark imposes — confirming that existing VLA policies are not ready for zero-tolerance lab settings. Applying the benchmark's RL post-training pipeline improves success rates by 37%, with the gains attributed to active error correction enabled by the interactive environment (recovering from execution drift rather than letting errors compound).
SafeLab targets a domain where the standard "failures are reversible" assumption breaks down entirely, making embodied safety a first-class objective rather than an afterthought. By pairing a high-fidelity chemistry-lab simulation with an LLM task generator, an automated expert, and an interactive RL loop, it provides not just an evaluation suite but a full pipeline for improving safety — and demonstrates that RL refinement substantially closes the gap that static imitation-learned VLAs leave open.
- ICML 2026: https://icml.cc/virtual/2026/poster/61584
← Back to ICML-2026