NeurIPS 2025 β VLA & Manipulation Survey
Compiled April 2026 (in retrospect). Focus: NeurIPS 2025 VLA / manipulation papers, grouped by approach, with a CoRL 2025 β NeurIPS 2025 β ICLR 2026 lineage map.
NeurIPS 2025 accepted 5,290 papers (24.52% from 21,575 submissions β a record 61% increase over 2024). No VLA / manipulation paper won an outstanding-paper award , but the field's presence via Spotlights and regular posters is the strongest yet. Six trends:
Dual-system VLAs become the default β Fast-in-Slow, ThinkAct, ChatVLA-2 attack reasoning-vs-action conflict with architectural separation.
Train-time + inference-time tricks reconcile VLMs with continuous action experts β Knowledge Insulation (Spotlight), Real-Time Chunking, CogVLA, Two-Steps Diffusion.
RL on pretrained VLAs is routine and characterized β PPO > DPO/GRPO (What-Can-RL-Bring), ReinFlow, Robot-R1, APO.
World models return as policy-conditioning signal β DreamVLA, VideoVLA, SAMPO, OSVI-WM, RLVR-World + EWM workshop.
Safety and failure become first-class β SafeVLA (Spotlight), SAFE, Latent Policy Barrier (Spotlight).
Humanoid + dex data scales via LLM-driven sim synthesis β HumanoidGen, DexFlyWheel, KungfuBot.
NeurIPS 2025 is the direct bridge between CoRL 2025 (Ο0.5, DSRL, DreamGen, Streaming Flow Policy) and ICLR 2026 (Ο0.6, Ο0.7, Embodied-R1, SimpleVLA-RL, Ctrl-World, Cosmos Policy). Knowledge Insulation is the clearest published hop β it's Ο0.5-KI formalized, and Ο0.6 / Ο0.7 adopt it directly.
Flagship: Knowledge Insulation
flowchart LR
V[Vision] --> VLM[VLM backbone]
L[Language] --> VLM
VLM -- FAST-token discrete actions --> CE[Cross-entropy loss<br/>VLM training]
VLM -- cond features --> AE[Continuous action expert]
AE -- flow matching --> A[Actions]
AE -. NO gradient back to VLM .-> VLM
note[Same VLM learns from<br/>discrete action tokens AND<br/>is conditioned by continuous head β<br/>but gradients are insulated]
Loading
Knowledge Insulation (Driess et al., PI β Spotlight ) is the published training formalism that turned Ο0.5 into Ο0.5-KI, which became the template for Ο0.6 and Ο0.7 . The key insight: use the VLM's own FAST-tokenized action targets to train its representations while letting a separate continuous action expert learn continuous control β but don't back-propagate the action-expert gradient into the VLM . That gradient insulation is what keeps the VLM's language-grounded knowledge intact while also enabling fast continuous control.
VLA Architecture & Dual-System
Training & Inference Efficiency
RL / Self-Improvement for VLAs
Diffusion / Flow-Matching Policies
World Models for Policy Learning
Safety / Robustness / Failure Detection
Dexterous / Humanoid / Mobile
Data & Benchmarks
Driving VLAs
Trends
CoRL 2025 β NeurIPS 2025 β ICLR 2026 Lineage
Reading List
1. VLA Architecture & Dual-System
Paper
Novelty
ChatVLA-2
Dynamic MoE that preserves VLM pretraining during action fine-tuning; beats OpenVLA / DexVLA / Ο0. 82.7% on open-world math-matching.
DreamVLA
Forecasts dynamic regions + depth + geometry + semantics + segmentation as action-forecasting signal. 76.7% real-robot success.
VLA-OS
Controlled ablation across OS-A / OS-I / OS-H paradigms. Finds visual-grounded > language planning; Hierarchical-VLA > Integrated / ActionOnly.
Fast-in-Slow
System-1 embedded inside System-2 via shared parameters (not cascaded). +8% sim / +11% real at 117.7 Hz.
ThinkAct (NVIDIA)
MLLM generates embodied reasoning plans; RL reward = goal completion + trajectory consistency. Plans compressed to visual latent for action head.
Chain-of-Action (ByteDance Seed)
Backward trajectory AR from a goal keyframe. SOTA on 60 RLBench + 8 real tasks.
BridgeVLA
Projects 3D inputs to multi-view 2D, unifies I/O as 2D heatmaps. RLBench 81.4 β 88.2%, real 96.8% on 10 tasks w/ 3 trajectories each.
CogVLA
FiLM-based instruction-driven encoder + LLM routing + token pruning. LIBERO 97.4%, 2.5Γ training / 2.8Γ inference speedup.
Hyper-GoalNet
HyperNetwork emits policy parameters from goal specs , not just input conditioning.
2. Training & Inference Efficiency
Paper
Novelty
Knowledge Insulation (Spotlight, PI)
Gradient-insulated VLM + action expert; published Ο0.5-KI recipe that seeds Ο0.6/Ο0.7.
Real-Time Chunking (PI + Berkeley)
Async inpainting of pending actions while previous chunk executes β works on any diffusion/flow VLA without retraining.
Two-Steps Diffusion Policy via genetic denoising
Population-based sampling stabilizes 2-step diffusion policies; 5 NFE competitive.
3. RL / Self-Improvement for VLAs
Paper
Novelty
What Can RL Bring to VLA Generalization?
Empirical: PPO > DPO / GRPO on VLA semantic + execution generalization.
ReinFlow
Injects learnable noise into deterministic flow path β Markov discrete-time process with exact likelihoods β enables online RL fine-tune of Ο0 / Ο0.5 / GR00T-N1.5. +135% reward on legged.
Robot-R1
DeepSeek-R1-style RL rewarding reasoning traces that lead to accurate keypoint predictions. 7B beats GPT-4o on low-level spatial reasoning.
APO (Action Preference Optimization)
Binary-signal RLHF-for-VLAs from on-deployment failures via HRI.
4. Diffusion / Flow-Matching Policies
Paper
Novelty
Real-Time Chunking
Async chunk inpainting for any diffusion/flow VLA.
ReinFlow
Online RL fine-tune via learned noise injection.
DynaGuide (Stanford β Shuran Song)
Latent dynamics model guides pretrained diffusion policies at denoising time toward/away from outcomes; up-weights rare behaviors.
AC-DiT (PKU)
Mobile-to-body conditioning + perception-aware 2D/3D fusion for mobile manipulation.
Two-Steps Diffusion Policy
Genetic denoising β 2-step diffusion competitive at 5 NFE.
5. World Models for Policy Learning
Paper
Novelty
DreamVLA
Multi-modal world-knowledge forecasting as inverse-dynamics signal.
VideoVLA
Multimodal DiT jointly predicts action chunks + future video; imagined futures correlate with task success.
SAMPO
Scale-wise intra-frame AR fused with temporal causal decoding for generative world models.
OSVI-WM
End-to-end learned world model for one-shot visual imitation on unseen tasks β no pretraining.
RLVR-World (THUML)
RL-with-verifiable-rewards for world-model quality (optimizes transition-prediction directly, not just likelihood).
6. Safety / Robustness / Failure Detection
This cluster is new at NeurIPS 2025 β VLA-specific safety had not formed a visible sub-literature at prior major conferences.
Paper
Novelty
SafeVLA (Spotlight, PKU-Alignment)
Constrained-learning safety alignment via CMDPs (Integrated Safety Approach, ISA). β83.58% violations, +3.85% task success. OOD-robust.
SAFE
Task-generic failure detector from VLA internal features; tested on OpenVLA, Ο0, Ο0-FAST. Conformal prediction trade-off.
Latent Policy Barrier (Spotlight, Stanford β Sun & Song)
Expert latent embeddings as implicit barrier; dynamics model optimizes future latents to stay on-manifold. Decouples imitation from OOD recovery.
7. Dexterous / Humanoid / Mobile Manipulation
Paper
Novelty
HumanoidGen (TeleHuman)
LLM composes atomic dex ops + relational constraints to auto-generate dual-arm dex tasks + demos.
DexFlyWheel
Self-improving cycle generates 2000+ dex demos across 4 tasks; 81.9% test β 78.3% real via digital twin.
KungfuBot (TeleHuman, Unitree G1)
Bi-level adaptive motion-tracking curriculum for highly-dynamic whole-body skills β deployed on Unitree G1.
Grasp2Grasp (Princeton)
Cross-morphology grasp transfer as stochastic transport via SchrΓΆdinger Bridges , with physics-informed costs.
AC-DiT (PKU)
Mobile-to-body conditioning for mobile manipulation diffusion policies.
OWMM-Agent
First dedicated foundation model for mobile manipulation ; agentic data-synthesis via VLM + PDDL + Habitat. Beats GPT-4o zero-shot.
CR-DAgger (Stanford β Song lab)
Compliance-mode DAgger + force-aware residual policy. +64% on contact-rich (book-flipping, belt, cable, gear).
Paper
Size / content
Robo2VLM-1 (Spotlight, Berkeley β Goldberg)
684,710 VQA from 463 scenes / 3,396 tasks / 176k real robot trajectories. Uses end-effector pose + gripper + force as non-visual ground-truth.
Impromptu VLA
80k curated corner-case driving clips (from 2M source) with planning-oriented QA. Improves NeuroNCAP closed-loop + nuScenes L2.
HumanoidGen
LLM-synthesized dex bimanual task suite
DexFlyWheel
2000+ dex demos, digital-twin sim-to-real
9. Driving VLAs (adjacent)
Paper
Novelty
AutoVLA (UCLA Mobility Lab)
Unified AR generative process with physical action tokens + CoT; dual fast/slow thinking + GRPO RFT. Top RFS-Spotlight on Waymo E2E.
Impromptu VLA
Structured corner-case taxonomy + QA-planning supervision.
Dual-system is the default architecture. Fast-in-Slow, ThinkAct, ChatVLA-2 all separate reasoning and action capacity via architectural splits.
VLM + action-expert reconciliation via gradient / inference tricks. Knowledge Insulation, Real-Time Chunking, CogVLA.
RL on pretrained VLAs is routine. PPO is the current winner (What-Can-RL-Bring); flow matching gets RL-ified (ReinFlow).
World models return as policy-conditioning signal. DreamVLA, VideoVLA, SAMPO, OSVI-WM, RLVR-World.
VLA safety becomes a first-class research cluster. SafeVLA, SAFE, Latent Policy Barrier β first major conference with this concentration.
Humanoid + dex data scales via LLM-driven sim synthesis. HumanoidGen, DexFlyWheel, KungfuBot.
11. CoRL 2025 β NeurIPS 2025 β ICLR 2026 Lineage
NeurIPS 2025 sits chronologically between CoRL 2025 (Sept) and ICLR 2026 (Apr). Direct published bridges:
CoRL 2025 β NeurIPS 2025 (backward lineage)
NeurIPS 2025 β ICLR 2026 (forward lineage)
NeurIPS 2025 seed
ICLR 2026 descendant
Knowledge Insulation
Ο0.6 β Ο*0.6 + RECAP β Ο0.7
ThinkAct / ChatVLA-2 / Fast-in-Slow
HAMLET , InstructVLA , WholeBodyVLA , HiMoE-VLA
Chain-of-Action + Robot-R1 + What-Can-RL-Bring
Embodied-R1 , SimpleVLA-RL , VLA-RFT , Stage-Aware RL , RL Tokens
DreamVLA + VideoVLA + SAMPO
Ctrl-World , Cosmos Policy , WorldGym
Knowledge Insulation + Real-Time Chunking + CogVLA
FASTER , Discrete Diffusion VLA , Unified Diffusion VLA , dVLA , DIVA & Fast-dVLA
VLA-OS
UniVLA , X-VLA , HyperVLA , VLM4VLA
Robo2VLM-1 + Impromptu VLA
RoboCasa365 , RoboArena β , EgoDex
CR-DAgger + APO
Hybrid Training , PLD
AC-DiT + OWMM-Agent
WholeBodyVLA , OmniSAT
SafeVLA + SAFE + Latent Policy Barrier
No 1:1 ICLR 2026 mirror yet β the safety sub-literature seeded here is still maturing. Likely follow-ups at ICLR 2027 / NeurIPS 2026.
Priority
Papers
Must read
Knowledge Insulation (Spotlight) , ChatVLA-2 , DreamVLA , ThinkAct
Spotlights
Knowledge Insulation , SafeVLA , Latent Policy Barrier , Robo2VLM-1
Architecture
ChatVLA-2 , DreamVLA , VLA-OS , Fast-in-Slow , Chain-of-Action
Training / Efficiency
Knowledge Insulation , Real-Time Chunking
RL for VLA
What-Can-RL-Bring , ReinFlow
World models
DreamVLA , VideoVLA
Safety
SafeVLA , Latent Policy Barrier
Data
Robo2VLM-1
Per-paper arXiv links appear on each paper's page.
HPT (Wang et al., arXiv 2409.20537) is NeurIPS 2024 , not 2025 β frequently miscited in third-party lists.
β Back to NeurIPS-2025 Β· NeurIPS Β· Home