NeurIPS 2025 VLA Manipulation Survey - Heungwoo/research GitHub Wiki

NeurIPS 2025 β€” VLA & Manipulation Survey

Compiled April 2026 (in retrospect). Focus: NeurIPS 2025 VLA / manipulation papers, grouped by approach, with a CoRL 2025 β†’ NeurIPS 2025 β†’ ICLR 2026 lineage map.

TL;DR

NeurIPS 2025 accepted 5,290 papers (24.52% from 21,575 submissions β€” a record 61% increase over 2024). No VLA / manipulation paper won an outstanding-paper award, but the field's presence via Spotlights and regular posters is the strongest yet. Six trends:

  1. Dual-system VLAs become the default β€” Fast-in-Slow, ThinkAct, ChatVLA-2 attack reasoning-vs-action conflict with architectural separation.
  2. Train-time + inference-time tricks reconcile VLMs with continuous action experts β€” Knowledge Insulation (Spotlight), Real-Time Chunking, CogVLA, Two-Steps Diffusion.
  3. RL on pretrained VLAs is routine and characterized β€” PPO > DPO/GRPO (What-Can-RL-Bring), ReinFlow, Robot-R1, APO.
  4. World models return as policy-conditioning signal β€” DreamVLA, VideoVLA, SAMPO, OSVI-WM, RLVR-World + EWM workshop.
  5. Safety and failure become first-class β€” SafeVLA (Spotlight), SAFE, Latent Policy Barrier (Spotlight).
  6. Humanoid + dex data scales via LLM-driven sim synthesis β€” HumanoidGen, DexFlyWheel, KungfuBot.

NeurIPS 2025 is the direct bridge between CoRL 2025 (Ο€0.5, DSRL, DreamGen, Streaming Flow Policy) and ICLR 2026 (Ο€0.6, Ο€0.7, Embodied-R1, SimpleVLA-RL, Ctrl-World, Cosmos Policy). Knowledge Insulation is the clearest published hop β€” it's Ο€0.5-KI formalized, and Ο€0.6 / Ο€0.7 adopt it directly.

Flagship: Knowledge Insulation

flowchart LR
  V[Vision] --> VLM[VLM backbone]
  L[Language] --> VLM
  VLM -- FAST-token discrete actions --> CE[Cross-entropy loss<br/>VLM training]
  VLM -- cond features --> AE[Continuous action expert]
  AE -- flow matching --> A[Actions]
  AE -. NO gradient back to VLM .-> VLM
  note[Same VLM learns from<br/>discrete action tokens AND<br/>is conditioned by continuous head β€”<br/>but gradients are insulated]
Loading

Knowledge Insulation (Driess et al., PI β€” Spotlight) is the published training formalism that turned Ο€0.5 into Ο€0.5-KI, which became the template for Ο€0.6 and Ο€0.7. The key insight: use the VLM's own FAST-tokenized action targets to train its representations while letting a separate continuous action expert learn continuous control β€” but don't back-propagate the action-expert gradient into the VLM. That gradient insulation is what keeps the VLM's language-grounded knowledge intact while also enabling fast continuous control.

Section index

  1. VLA Architecture & Dual-System
  2. Training & Inference Efficiency
  3. RL / Self-Improvement for VLAs
  4. Diffusion / Flow-Matching Policies
  5. World Models for Policy Learning
  6. Safety / Robustness / Failure Detection
  7. Dexterous / Humanoid / Mobile
  8. Data & Benchmarks
  9. Driving VLAs
  10. Trends
  11. CoRL 2025 β†’ NeurIPS 2025 β†’ ICLR 2026 Lineage
  12. Reading List

1. VLA Architecture & Dual-System

Paper Novelty
ChatVLA-2 Dynamic MoE that preserves VLM pretraining during action fine-tuning; beats OpenVLA / DexVLA / Ο€0. 82.7% on open-world math-matching.
DreamVLA Forecasts dynamic regions + depth + geometry + semantics + segmentation as action-forecasting signal. 76.7% real-robot success.
VLA-OS Controlled ablation across OS-A / OS-I / OS-H paradigms. Finds visual-grounded > language planning; Hierarchical-VLA > Integrated / ActionOnly.
Fast-in-Slow System-1 embedded inside System-2 via shared parameters (not cascaded). +8% sim / +11% real at 117.7 Hz.
ThinkAct (NVIDIA) MLLM generates embodied reasoning plans; RL reward = goal completion + trajectory consistency. Plans compressed to visual latent for action head.
Chain-of-Action (ByteDance Seed) Backward trajectory AR from a goal keyframe. SOTA on 60 RLBench + 8 real tasks.
BridgeVLA Projects 3D inputs to multi-view 2D, unifies I/O as 2D heatmaps. RLBench 81.4 β†’ 88.2%, real 96.8% on 10 tasks w/ 3 trajectories each.
CogVLA FiLM-based instruction-driven encoder + LLM routing + token pruning. LIBERO 97.4%, 2.5Γ— training / 2.8Γ— inference speedup.
Hyper-GoalNet HyperNetwork emits policy parameters from goal specs, not just input conditioning.

2. Training & Inference Efficiency

Paper Novelty
Knowledge Insulation (Spotlight, PI) Gradient-insulated VLM + action expert; published Ο€0.5-KI recipe that seeds Ο€0.6/Ο€0.7.
Real-Time Chunking (PI + Berkeley) Async inpainting of pending actions while previous chunk executes β€” works on any diffusion/flow VLA without retraining.
Two-Steps Diffusion Policy via genetic denoising Population-based sampling stabilizes 2-step diffusion policies; 5 NFE competitive.

3. RL / Self-Improvement for VLAs

Paper Novelty
What Can RL Bring to VLA Generalization? Empirical: PPO > DPO / GRPO on VLA semantic + execution generalization.
ReinFlow Injects learnable noise into deterministic flow path β†’ Markov discrete-time process with exact likelihoods β†’ enables online RL fine-tune of Ο€0 / Ο€0.5 / GR00T-N1.5. +135% reward on legged.
Robot-R1 DeepSeek-R1-style RL rewarding reasoning traces that lead to accurate keypoint predictions. 7B beats GPT-4o on low-level spatial reasoning.
APO (Action Preference Optimization) Binary-signal RLHF-for-VLAs from on-deployment failures via HRI.

4. Diffusion / Flow-Matching Policies

Paper Novelty
Real-Time Chunking Async chunk inpainting for any diffusion/flow VLA.
ReinFlow Online RL fine-tune via learned noise injection.
DynaGuide (Stanford β€” Shuran Song) Latent dynamics model guides pretrained diffusion policies at denoising time toward/away from outcomes; up-weights rare behaviors.
AC-DiT (PKU) Mobile-to-body conditioning + perception-aware 2D/3D fusion for mobile manipulation.
Two-Steps Diffusion Policy Genetic denoising β†’ 2-step diffusion competitive at 5 NFE.

5. World Models for Policy Learning

Paper Novelty
DreamVLA Multi-modal world-knowledge forecasting as inverse-dynamics signal.
VideoVLA Multimodal DiT jointly predicts action chunks + future video; imagined futures correlate with task success.
SAMPO Scale-wise intra-frame AR fused with temporal causal decoding for generative world models.
OSVI-WM End-to-end learned world model for one-shot visual imitation on unseen tasks β€” no pretraining.
RLVR-World (THUML) RL-with-verifiable-rewards for world-model quality (optimizes transition-prediction directly, not just likelihood).

6. Safety / Robustness / Failure Detection

This cluster is new at NeurIPS 2025 β€” VLA-specific safety had not formed a visible sub-literature at prior major conferences.

Paper Novelty
SafeVLA (Spotlight, PKU-Alignment) Constrained-learning safety alignment via CMDPs (Integrated Safety Approach, ISA). βˆ’83.58% violations, +3.85% task success. OOD-robust.
SAFE Task-generic failure detector from VLA internal features; tested on OpenVLA, Ο€0, Ο€0-FAST. Conformal prediction trade-off.
Latent Policy Barrier (Spotlight, Stanford β€” Sun & Song) Expert latent embeddings as implicit barrier; dynamics model optimizes future latents to stay on-manifold. Decouples imitation from OOD recovery.

7. Dexterous / Humanoid / Mobile Manipulation

Paper Novelty
HumanoidGen (TeleHuman) LLM composes atomic dex ops + relational constraints to auto-generate dual-arm dex tasks + demos.
DexFlyWheel Self-improving cycle generates 2000+ dex demos across 4 tasks; 81.9% test β†’ 78.3% real via digital twin.
KungfuBot (TeleHuman, Unitree G1) Bi-level adaptive motion-tracking curriculum for highly-dynamic whole-body skills β€” deployed on Unitree G1.
Grasp2Grasp (Princeton) Cross-morphology grasp transfer as stochastic transport via SchrΓΆdinger Bridges, with physics-informed costs.
AC-DiT (PKU) Mobile-to-body conditioning for mobile manipulation diffusion policies.
OWMM-Agent First dedicated foundation model for mobile manipulation; agentic data-synthesis via VLM + PDDL + Habitat. Beats GPT-4o zero-shot.
CR-DAgger (Stanford β€” Song lab) Compliance-mode DAgger + force-aware residual policy. +64% on contact-rich (book-flipping, belt, cable, gear).

8. Data & Benchmarks

Paper Size / content
Robo2VLM-1 (Spotlight, Berkeley β€” Goldberg) 684,710 VQA from 463 scenes / 3,396 tasks / 176k real robot trajectories. Uses end-effector pose + gripper + force as non-visual ground-truth.
Impromptu VLA 80k curated corner-case driving clips (from 2M source) with planning-oriented QA. Improves NeuroNCAP closed-loop + nuScenes L2.
HumanoidGen LLM-synthesized dex bimanual task suite
DexFlyWheel 2000+ dex demos, digital-twin sim-to-real

9. Driving VLAs (adjacent)

Paper Novelty
AutoVLA (UCLA Mobility Lab) Unified AR generative process with physical action tokens + CoT; dual fast/slow thinking + GRPO RFT. Top RFS-Spotlight on Waymo E2E.
Impromptu VLA Structured corner-case taxonomy + QA-planning supervision.

10. Trends

  1. Dual-system is the default architecture. Fast-in-Slow, ThinkAct, ChatVLA-2 all separate reasoning and action capacity via architectural splits.
  2. VLM + action-expert reconciliation via gradient / inference tricks. Knowledge Insulation, Real-Time Chunking, CogVLA.
  3. RL on pretrained VLAs is routine. PPO is the current winner (What-Can-RL-Bring); flow matching gets RL-ified (ReinFlow).
  4. World models return as policy-conditioning signal. DreamVLA, VideoVLA, SAMPO, OSVI-WM, RLVR-World.
  5. VLA safety becomes a first-class research cluster. SafeVLA, SAFE, Latent Policy Barrier β€” first major conference with this concentration.
  6. Humanoid + dex data scales via LLM-driven sim synthesis. HumanoidGen, DexFlyWheel, KungfuBot.

11. CoRL 2025 β†’ NeurIPS 2025 β†’ ICLR 2026 Lineage

NeurIPS 2025 sits chronologically between CoRL 2025 (Sept) and ICLR 2026 (Apr). Direct published bridges:

CoRL 2025 β†’ NeurIPS 2025 (backward lineage)

CoRL 2025 seed NeurIPS 2025 descendant
Ο€0.5 (Oral, co-training + hierarchical) Knowledge Insulation (Spotlight) β€” same recipe, formalized
DSRL (RL in noise-latent space) ReinFlow β€” extends to flow matching + Ο€0 / Ο€0.5 / GR00T targets
Streaming Flow Policy Real-Time Chunking β€” async chunk inpainting variant
ECoT-Lite Chain-of-Action + Robot-R1 + ThinkAct β€” full CoT for actions
DreamGen DreamVLA + VideoVLA β€” video-as-supervision matures
RoboMonkey (test-time scaling) SAFE β€” failure detection on the same OpenVLA / Ο€0 / Ο€0-FAST targets
X-Sim DexFlyWheel β€” sim-to-real self-improving variant
ManipBench Robo2VLM-1 (Spotlight) β€” larger VQA dataset from real robot trajectories

NeurIPS 2025 β†’ ICLR 2026 (forward lineage)

NeurIPS 2025 seed ICLR 2026 descendant
Knowledge Insulation Ο€0.6 β†’ Ο€*0.6 + RECAP β†’ Ο€0.7
ThinkAct / ChatVLA-2 / Fast-in-Slow HAMLET, InstructVLA, WholeBodyVLA, HiMoE-VLA
Chain-of-Action + Robot-R1 + What-Can-RL-Bring Embodied-R1, SimpleVLA-RL, VLA-RFT, Stage-Aware RL, RL Tokens
DreamVLA + VideoVLA + SAMPO Ctrl-World, Cosmos Policy, WorldGym
Knowledge Insulation + Real-Time Chunking + CogVLA FASTER, Discrete Diffusion VLA, Unified Diffusion VLA, dVLA, DIVA & Fast-dVLA
VLA-OS UniVLA, X-VLA, HyperVLA, VLM4VLA
Robo2VLM-1 + Impromptu VLA RoboCasa365, RoboArena ∞, EgoDex
CR-DAgger + APO Hybrid Training, PLD
AC-DiT + OWMM-Agent WholeBodyVLA, OmniSAT
SafeVLA + SAFE + Latent Policy Barrier No 1:1 ICLR 2026 mirror yet β€” the safety sub-literature seeded here is still maturing. Likely follow-ups at ICLR 2027 / NeurIPS 2026.

12. Reading List

Priority Papers
Must read Knowledge Insulation (Spotlight), ChatVLA-2, DreamVLA, ThinkAct
Spotlights Knowledge Insulation, SafeVLA, Latent Policy Barrier, Robo2VLM-1
Architecture ChatVLA-2, DreamVLA, VLA-OS, Fast-in-Slow, Chain-of-Action
Training / Efficiency Knowledge Insulation, Real-Time Chunking
RL for VLA What-Can-RL-Bring, ReinFlow
World models DreamVLA, VideoVLA
Safety SafeVLA, Latent Policy Barrier
Data Robo2VLM-1

Sources

Per-paper arXiv links appear on each paper's page.

Correction

HPT (Wang et al., arXiv 2409.20537) is NeurIPS 2024, not 2025 β€” frequently miscited in third-party lists.

← Back to NeurIPS-2025 Β· NeurIPS Β· Home

⚠️ **GitHub.com Fallback** ⚠️