NeurIPS 2025 - Heungwoo/research GitHub Wiki
NeurIPS 2025
Neural Information Processing Systems 2025 β San Diego (primary) + Mexico City (secondary), Dec 2β7, 2025.
Scale: 21,575 submissions β 5,290 accepted (24.52%). Record 61% submission increase over 2024. Two physical sites; new Journal Track + Position Paper Track.
No VLA / manipulation paper won an outstanding-paper award this year β the Physical-AI cluster landed in spotlights and regular posters rather than award slots.
Surveys hosted in this wiki
- VLA & Manipulation Survey (NeurIPS 2025) β 33 confirmed VLA/manipulation papers, 6 categories, 6 trends, and CoRL 2025 β NeurIPS 2025 β ICLR 2026 lineage map.
Spotlights relevant to VLA / manipulation
| Paper | What it is |
|---|---|
| [Knowledge Insulation](/Heungwoo/research/wiki/NeurIPS-2025-Knowledge-Insulation) (Spotlight) | Physical Intelligence's formalization of the training recipe used in Ο0.5-KI β Ο0.6 β Ο0.7. Gradient-insulates the VLM from the action expert. |
| [SafeVLA](/Heungwoo/research/wiki/NeurIPS-2025-SafeVLA) (Spotlight) | First constrained-learning safety alignment for VLAs. β83.58% safety violations, +3.85% task success. |
| [Latent Policy Barrier](/Heungwoo/research/wiki/NeurIPS-2025-Latent-Policy-Barrier) (Spotlight) | Treats expert latent embeddings as an implicit barrier; decouples imitation from OOD recovery. |
| [Robo2VLM-1](/Heungwoo/research/wiki/NeurIPS-2025-Robo2VLM-1) (Spotlight) | 684,710 VQA from 463 scenes / 176k real trajectories. Uses robot proprioception as non-visual ground-truth. |
Quick paper index (by category)
VLA architecture (backbones / action heads / CoT)
- ChatVLA-2 β Dynamic MoE preserving VLM pretraining
- DreamVLA β World-knowledge-conditioned VLA (multi-modal forecasting)
- VLA-OS β Controlled architecture ablation suite (OS-A / OS-I / OS-H)
- Chain-of-Action β Backward trajectory AR from goal keyframe
- BridgeVLA β 2D heatmap I/O unification for 3D manipulation (arXiv 2506.07961)
- CogVLA β Instruction-driven token sparsification, 2.5Γ training / 2.8Γ inference speedup (arXiv 2508.21046)
Dual-system (System 1 + System 2)
- Fast-in-Slow β Embedded (not cascaded) S1-in-S2 via shared parameters
- ThinkAct (NVIDIA) β Reinforced visual latent plans
Training / inference efficiency
- Knowledge Insulation (Spotlight, PI) β Gradient-insulated VLM + action expert
- Real-Time Chunking (PI + Berkeley) β Async inpainting over action chunks
- Two-Steps Diffusion Policy via genetic denoising β 2-step diffusion competitive at 5 NFE
RL / self-improvement for VLAs
- What Can RL Bring to VLA Generalization? β PPO > DPO/GRPO empirically
- ReinFlow β Online RL fine-tuning of flow-matching policies (Ο0/Ο0.5/GR00T-N1.5)
- Robot-R1 β R1-style RL for reasoning-grounded keypoint prediction (arXiv 2506.00070)
- APO (Action Preference Optimization) β Binary-preference RL from human interventions (arXiv 2506.07127)
Diffusion / flow-matching policies
- DynaGuide (Stanford) β Dynamics-model-guided inference-time steering (arXiv 2506.13922)
- AC-DiT (PKU) β Mobile-to-body conditioning for mobile manipulation (arXiv 2507.01961)
- Real-Time Chunking
- ReinFlow
World models for policy learning
- DreamVLA β Depth / geometry / segmentation / dynamics forecasting
- VideoVLA β Joint action chunk + future-video prediction
- SAMPO β Scale-wise autoregressive world model (arXiv 2509.15536)
- OSVI-WM β End-to-end world model for one-shot visual imitation (arXiv 2505.20425)
- RLVR-World (THUML) β RL with verifiable rewards for world-model quality
Safety / robustness / alignment
- SafeVLA (Spotlight)
- SAFE β Task-generic failure detection from VLA internals (tested on OpenVLA, Ο0, Ο0-FAST)
- Latent Policy Barrier (Spotlight)
Human-in-the-loop / preference / intervention
- CR-DAgger (Stanford) β Compliance-mode DAgger + force-aware residual, +64% on contact-rich (arXiv 2506.16685)
- APO β Human-assisted preference optimization
Dexterous / bimanual / humanoid
- HumanoidGen β LLM-driven dexterous bimanual task synthesis
- DexFlyWheel β Self-improving data flywheel for dex sim-to-real (arXiv 2509.23829)
- KungfuBot β Dynamic whole-body control on Unitree G1 (arXiv 2506.12851)
- Grasp2Grasp β SchrΓΆdinger-Bridge cross-embodiment grasp translation (arXiv 2506.02489)
Mobile manipulation
- AC-DiT β (arXiv 2507.01961)
- OWMM-Agent β Agentic data-synthesis foundation model for mobile manipulation (arXiv 2506.04217)
Data & benchmarks
- Robo2VLM-1 (Spotlight, Berkeley) β 684,710 VQA from robot trajectories
- Impromptu VLA β 80k curated corner-case driving clips (arXiv 2505.23757)
- HumanoidGen β Dex bimanual auto-generated suite
- DexFlyWheel β 2000+ dex demos, 78.3% real transfer
Driving VLAs
- AutoVLA (UCLA) β Adaptive-reasoning GRPO + Waymo E2E challenge (arXiv 2506.13757)
- Impromptu VLA
Relevant NeurIPS 2025 workshops
| Workshop | Focus |
|---|---|
| SpaVLE | Space in Vision, Language, and Embodied AI (hosts MARS multi-agent embodied challenge) |
| EWM | Embodied World Models for Decision Making |
| LAW | Bridging Language, Agent, and World Models |
| 7th Robot Learning Workshop | "Towards Robots with Human-Level Abilities" |
| EAI Challenge | Embodied Agent Interface Challenge |
| Scaling Environments for Agents | Scalable sim environments for LLM agents |
| Embodied and Safe-Assured Robotic Systems (CDMX) | Safety/QA for autonomous robots |
Trends (6)
- Dual-system VLAs become the default recipe. Fast-in-Slow, ThinkAct, ChatVLA-2 all address reasoning-vs-action conflict with architectural separation of concerns.
- Train-time + inference-time tricks reconcile VLMs with continuous action experts. Knowledge Insulation, Real-Time Chunking, CogVLA, Two-Steps Diffusion Policy.
- RL on pretrained VLAs is routine and empirically characterized. PPO > DPO/GRPO (What-Can-RL-Bring), ReinFlow, Robot-R1, APO, AutoVLA.
- World models return as policy-conditioning signal. DreamVLA, VideoVLA, SAMPO, OSVI-WM, RLVR-World + the EWM workshop.
- Safety and failure become first-class. SafeVLA (Spotlight), SAFE, Latent Policy Barrier (Spotlight) β first major conference where VLA-specific safety forms a visible cluster.
- Humanoid + dex data scales via LLM-driven sim synthesis. HumanoidGen, DexFlyWheel, KungfuBot.
Correction
HPT (Wang et al., arXiv 2409.20537) is NeurIPS 2024, not 2025 β frequently miscited. Correct venue: https://neurips.cc/virtual/2024/poster/95294.