NeurIPS 2025 Latent Policy Barrier - Heungwoo/research GitHub Wiki
Venue: NeurIPS 2025 (Spotlight) ยท Authors: Zhanyi Sun & Shuran Song (Stanford, REAL Lab) ยท arXiv: 2508.05941 Category: Safety / Robustness / Diffusion
flowchart LR
Obs[Observation] --> E[Policy encoder]
E --> L[Latent z]
L --> ExpertMani["Expert latent manifold<br/>(implicit from demos)"]
E -- OOD? --> B{barrier:<br/>stay on manifold}
B -- dynamics model --> Future[Optimize future latents<br/>to return to manifold]
Future --> A[Action]
note[Treats expert-trajectory latents as<br/>an implicit control barrier function]
Imitation-learning policies work on in-distribution states but fail catastrophically OOD โ small perturbations push the policy outside the expert manifold where it has no information. Standard fixes (data augmentation, ensembles, RL fine-tuning) improve average performance but don't give a principled "I'm outside my safe region" signal.
Use expert latent embeddings as an implicit control barrier:
- Train a base diffusion policy on expert data only, and separately train a latent dynamics model on a mix of expert + suboptimal policy rollout data (unlabeled โ no rewards, no human corrections).
- The expert-demonstration latents define an implicit barrier/manifold in latent space.
- At inference, the dynamics model predicts future latents and steers the policy in latent space by minimizing the distance between predicted future latents and their nearest neighbors among expert-demonstration latents โ pulling execution back in-distribution.
- Formally inspired by Control Barrier Functions: treat distance-from-expert-manifold as the barrier.
This decouples imitation (primary task) from OOD recovery (safety) โ the two can be trained separately and combined at inference.
Baselines: Expert BC (diffusion policy on expert data only), Mixed BC, Filtered BC, CQL (offline RL), and CCIL.
- Simulation (limited-demo, 20% of full data): LPB matches or exceeds every baseline on all four simulated tasks. Reported success rates include Transport 0.85 (vs. 0.68 Expert BC), Tool-Hang 0.39 (vs. 0.27 Expert BC), Square 0.65.
- Real robot: Belt assembly (NIST board) improves to 0.75 success vs. 0.55 for the base policy; cup arrangement substantially outperforms the base policy on out-of-distribution initial poses while matching it on in-distribution poses.
- Maintains on-distribution performance (doesn't over-conservatively block normal actions).
- Augments a pretrained diffusion policy at inference time without retraining it.
NeurIPS 2025 Spotlight. Along with SafeVLA (training-time safety) and SAFE (detection-time), LPB is the inference-time pillar of VLA safety that emerges at NeurIPS 2025.
Bridges two literatures:
- From safe control theory: Control Barrier Functions have been used for decades in robotics; LPB puts them in latent space and learns them from data.
- From diffusion policy: Fits with Stanford's DynaGuide (NeurIPS 2025) โ both steer diffusion policies at inference using a learned dynamics model, but with different objectives (guidance vs. barrier).
- arXiv: https://arxiv.org/abs/2508.05941
- NeurIPS virtual page: https://neurips.cc/virtual/2025/loc/san-diego/poster/119037
- SafeVLA (training-time safety sibling)
- Survey: VLA & Manipulation (ICLR 2026)
โ Back to NeurIPS-2025