NeurIPS 2025 - Heungwoo/research GitHub Wiki

NeurIPS 2025

Neural Information Processing Systems 2025 β€” San Diego (primary) + Mexico City (secondary), Dec 2–7, 2025.

Scale: 21,575 submissions β†’ 5,290 accepted (24.52%). Record 61% submission increase over 2024. Two physical sites; new Journal Track + Position Paper Track.

No VLA / manipulation paper won an outstanding-paper award this year β€” the Physical-AI cluster landed in spotlights and regular posters rather than award slots.

Surveys hosted in this wiki

Spotlights relevant to VLA / manipulation

Paper What it is
[Knowledge Insulation](/Heungwoo/research/wiki/NeurIPS-2025-Knowledge-Insulation) (Spotlight) Physical Intelligence's formalization of the training recipe used in Ο€0.5-KI β†’ Ο€0.6 β†’ Ο€0.7. Gradient-insulates the VLM from the action expert.
[SafeVLA](/Heungwoo/research/wiki/NeurIPS-2025-SafeVLA) (Spotlight) First constrained-learning safety alignment for VLAs. βˆ’83.58% safety violations, +3.85% task success.
[Latent Policy Barrier](/Heungwoo/research/wiki/NeurIPS-2025-Latent-Policy-Barrier) (Spotlight) Treats expert latent embeddings as an implicit barrier; decouples imitation from OOD recovery.
[Robo2VLM-1](/Heungwoo/research/wiki/NeurIPS-2025-Robo2VLM-1) (Spotlight) 684,710 VQA from 463 scenes / 176k real trajectories. Uses robot proprioception as non-visual ground-truth.

Quick paper index (by category)

VLA architecture (backbones / action heads / CoT)

  • ChatVLA-2 β€” Dynamic MoE preserving VLM pretraining
  • DreamVLA β€” World-knowledge-conditioned VLA (multi-modal forecasting)
  • VLA-OS β€” Controlled architecture ablation suite (OS-A / OS-I / OS-H)
  • Chain-of-Action β€” Backward trajectory AR from goal keyframe
  • BridgeVLA β€” 2D heatmap I/O unification for 3D manipulation (arXiv 2506.07961)
  • CogVLA β€” Instruction-driven token sparsification, 2.5Γ— training / 2.8Γ— inference speedup (arXiv 2508.21046)

Dual-system (System 1 + System 2)

  • Fast-in-Slow β€” Embedded (not cascaded) S1-in-S2 via shared parameters
  • ThinkAct (NVIDIA) β€” Reinforced visual latent plans

Training / inference efficiency

  • Knowledge Insulation (Spotlight, PI) β€” Gradient-insulated VLM + action expert
  • Real-Time Chunking (PI + Berkeley) β€” Async inpainting over action chunks
  • Two-Steps Diffusion Policy via genetic denoising β€” 2-step diffusion competitive at 5 NFE

RL / self-improvement for VLAs

Diffusion / flow-matching policies

World models for policy learning

  • DreamVLA β€” Depth / geometry / segmentation / dynamics forecasting
  • VideoVLA β€” Joint action chunk + future-video prediction
  • SAMPO β€” Scale-wise autoregressive world model (arXiv 2509.15536)
  • OSVI-WM β€” End-to-end world model for one-shot visual imitation (arXiv 2505.20425)
  • RLVR-World (THUML) β€” RL with verifiable rewards for world-model quality

Safety / robustness / alignment

  • SafeVLA (Spotlight)
  • SAFE β€” Task-generic failure detection from VLA internals (tested on OpenVLA, Ο€0, Ο€0-FAST)
  • Latent Policy Barrier (Spotlight)

Human-in-the-loop / preference / intervention

  • CR-DAgger (Stanford) β€” Compliance-mode DAgger + force-aware residual, +64% on contact-rich (arXiv 2506.16685)
  • APO β€” Human-assisted preference optimization

Dexterous / bimanual / humanoid

  • HumanoidGen β€” LLM-driven dexterous bimanual task synthesis
  • DexFlyWheel β€” Self-improving data flywheel for dex sim-to-real (arXiv 2509.23829)
  • KungfuBot β€” Dynamic whole-body control on Unitree G1 (arXiv 2506.12851)
  • Grasp2Grasp β€” SchrΓΆdinger-Bridge cross-embodiment grasp translation (arXiv 2506.02489)

Mobile manipulation

Data & benchmarks

  • Robo2VLM-1 (Spotlight, Berkeley) β€” 684,710 VQA from robot trajectories
  • Impromptu VLA β€” 80k curated corner-case driving clips (arXiv 2505.23757)
  • HumanoidGen β€” Dex bimanual auto-generated suite
  • DexFlyWheel β€” 2000+ dex demos, 78.3% real transfer

Driving VLAs

  • AutoVLA (UCLA) β€” Adaptive-reasoning GRPO + Waymo E2E challenge (arXiv 2506.13757)
  • Impromptu VLA

Relevant NeurIPS 2025 workshops

Workshop Focus
SpaVLE Space in Vision, Language, and Embodied AI (hosts MARS multi-agent embodied challenge)
EWM Embodied World Models for Decision Making
LAW Bridging Language, Agent, and World Models
7th Robot Learning Workshop "Towards Robots with Human-Level Abilities"
EAI Challenge Embodied Agent Interface Challenge
Scaling Environments for Agents Scalable sim environments for LLM agents
Embodied and Safe-Assured Robotic Systems (CDMX) Safety/QA for autonomous robots

Trends (6)

  1. Dual-system VLAs become the default recipe. Fast-in-Slow, ThinkAct, ChatVLA-2 all address reasoning-vs-action conflict with architectural separation of concerns.
  2. Train-time + inference-time tricks reconcile VLMs with continuous action experts. Knowledge Insulation, Real-Time Chunking, CogVLA, Two-Steps Diffusion Policy.
  3. RL on pretrained VLAs is routine and empirically characterized. PPO > DPO/GRPO (What-Can-RL-Bring), ReinFlow, Robot-R1, APO, AutoVLA.
  4. World models return as policy-conditioning signal. DreamVLA, VideoVLA, SAMPO, OSVI-WM, RLVR-World + the EWM workshop.
  5. Safety and failure become first-class. SafeVLA (Spotlight), SAFE, Latent Policy Barrier (Spotlight) β€” first major conference where VLA-specific safety forms a visible cluster.
  6. Humanoid + dex data scales via LLM-driven sim synthesis. HumanoidGen, DexFlyWheel, KungfuBot.

Correction

HPT (Wang et al., arXiv 2409.20537) is NeurIPS 2024, not 2025 β€” frequently miscited. Correct venue: https://neurips.cc/virtual/2024/poster/95294.

← Back to NeurIPS Β· Home