NeurIPS 2025 SafeVLA - Heungwoo/research GitHub Wiki
Venue: NeurIPS 2025 (Spotlight) ยท Authors: Borong Zhang, Yuhao Zhang, Jiaming Ji, Yingshan Lei, Yishuai Cai, Josef Dai, Yuanpei Chen, Yaodong Yang (PKU-Alignment, Peking University) ยท arXiv: 2503.03480 Category: Safety / Robustness
flowchart LR
Base[VLA training] --> CL[Constrained learning:<br/>CMDP + Integrated Safety Approach ISA]
CL -- reward-first --> S[Reward signal: task success]
CL -- constraints --> Sf[Safety signal: violation cost]
S & Sf --> Balance[Solve constrained optimization]
Balance --> VLA[SafeVLA policy]
VLA -- eval --> R[-83.58% violations<br/>+3.85% success<br/>OOD robust]
VLAs trained by plain imitation learning occasionally take unsafe actions โ collisions, excessive force, tipping, spills. There's no principled mechanism in the training objective to penalize these; they come out as rare but costly failures. LLM-alignment tools (RLHF, DPO) don't directly translate because embodied safety is about physical consequences, not text preferences.
Formulate VLA training as a Constrained Markov Decision Process (CMDP):
- Primary objective: maximize task success.
- Constraints: safety violation cost โค threshold.
- Solve using the Integrated Safety Approach (ISA) โ a pipeline that models safety requirements, actively elicits diverse unsafe behaviors, constrains the policy via safe RL (CMDP with Lagrangian / min-max optimization against elicited risks), and assures safety through targeted evaluation.
Constraints cover an Object Safety Constraint (penalizing unintended object displacement/rotation) and a Robot Safety Constraint (preventing collisions with forbidden structures), observable during rollouts.
Built on the SPOC transformer VLA, fine-tuned and evaluated in Safety-CHORES โ a new AI2-THOR / ProcTHOR-based benchmark with millions of unique scenes that extends the CHORES task suite with safety constraints.
- โ83.58% cumulative safety-violation cost over SOTA baselines.
- +3.85% task success rate โ safety doesn't have to cost performance.
- Generalizes to OOD perturbations (unseen objects, lighting, initial configurations).
First algorithm to explicitly incorporate safety constraints into VLAs (and first comprehensive VLA safety benchmark, Safety-CHORES) โ seeds an emerging sub-literature at NeurIPS 2025 alongside SAFE (failure detection) and Latent Policy Barrier (OOD recovery).
Before NeurIPS 2025, VLA safety was scattered. After it, three complementary approaches form a proto-stack:
- Training-time: SafeVLA (CMDP)
- Detection-time: SAFE (failure classifier)
- Inference-time: Latent Policy Barrier (stay on expert manifold)
No direct 1:1 ICLR 2026 descendant yet โ safety is still maturing.
- arXiv: https://arxiv.org/abs/2503.03480
- NeurIPS virtual page: https://neurips.cc/virtual/2025/poster/116975
- GitHub: https://github.com/PKU-Alignment/SafeVLA
- Latent Policy Barrier (inference-time safety sibling)
- Survey: VLA & Manipulation (ICLR 2026)
โ Back to NeurIPS-2025