NeurIPS 2025 SafeVLA - Heungwoo/research GitHub Wiki

SafeVLA โ€” Safety Alignment via Constrained Learning

Venue: NeurIPS 2025 (Spotlight) ยท Authors: Borong Zhang, Yuhao Zhang, Jiaming Ji, Yingshan Lei, Yishuai Cai, Josef Dai, Yuanpei Chen, Yaodong Yang (PKU-Alignment, Peking University) ยท arXiv: 2503.03480 Category: Safety / Robustness

Approach diagram

flowchart LR
  Base[VLA training] --> CL[Constrained learning:<br/>CMDP + Integrated Safety Approach ISA]
  CL -- reward-first --> S[Reward signal: task success]
  CL -- constraints --> Sf[Safety signal: violation cost]
  S & Sf --> Balance[Solve constrained optimization]
  Balance --> VLA[SafeVLA policy]
  VLA -- eval --> R[-83.58% violations<br/>+3.85% success<br/>OOD robust]
Loading

Problem

VLAs trained by plain imitation learning occasionally take unsafe actions โ€” collisions, excessive force, tipping, spills. There's no principled mechanism in the training objective to penalize these; they come out as rare but costly failures. LLM-alignment tools (RLHF, DPO) don't directly translate because embodied safety is about physical consequences, not text preferences.

Method

Formulate VLA training as a Constrained Markov Decision Process (CMDP):

  • Primary objective: maximize task success.
  • Constraints: safety violation cost โ‰ค threshold.
  • Solve using the Integrated Safety Approach (ISA) โ€” a pipeline that models safety requirements, actively elicits diverse unsafe behaviors, constrains the policy via safe RL (CMDP with Lagrangian / min-max optimization against elicited risks), and assures safety through targeted evaluation.

Constraints cover an Object Safety Constraint (penalizing unintended object displacement/rotation) and a Robot Safety Constraint (preventing collisions with forbidden structures), observable during rollouts.

Built on the SPOC transformer VLA, fine-tuned and evaluated in Safety-CHORES โ€” a new AI2-THOR / ProcTHOR-based benchmark with millions of unique scenes that extends the CHORES task suite with safety constraints.

Results

  • โˆ’83.58% cumulative safety-violation cost over SOTA baselines.
  • +3.85% task success rate โ€” safety doesn't have to cost performance.
  • Generalizes to OOD perturbations (unseen objects, lighting, initial configurations).

Significance

First algorithm to explicitly incorporate safety constraints into VLAs (and first comprehensive VLA safety benchmark, Safety-CHORES) โ€” seeds an emerging sub-literature at NeurIPS 2025 alongside SAFE (failure detection) and Latent Policy Barrier (OOD recovery).

Before NeurIPS 2025, VLA safety was scattered. After it, three complementary approaches form a proto-stack:

  • Training-time: SafeVLA (CMDP)
  • Detection-time: SAFE (failure classifier)
  • Inference-time: Latent Policy Barrier (stay on expert manifold)

No direct 1:1 ICLR 2026 descendant yet โ€” safety is still maturing.

Links

Related pages

โ† Back to NeurIPS-2025

โš ๏ธ **GitHub.com Fallback** โš ๏ธ