CoRL 2025 VLA Manipulation Survey - Heungwoo/research GitHub Wiki
Compiled April 2026 (in retrospect). Focus: VLA architecture, training & inference recipes, dexterous / humanoid manipulation, cross-embodiment, diffusion & flow policies, world models, benchmarks.
CoRL 2025 (Seoul, Sept 27โ30) accepted 221 papers out of 263 submissions (84.03% acceptance), of which 42 were orals. From the VLA/manipulation perspective, six trends stand out:
- Hierarchical VLAs beat monolithic ones for open-world generalization โ exemplified by ฯ0.5 (CoRL 2025 Oral + Best Paper Award Finalist), which sets up the entire ฯ series that the field now benchmarks against.
- Test-time scaling + trajectory streaming โ RoboMonkey (best-of-N with verifier), Streaming Flow Policy, DemoSpeedup, SAIL.
- Human video as the default cross-embodiment data source โ DexUMI, Visual Imitation โ Humanoid, UniSkill, ImMimic, X-Sim.
- Diffusion / flow policies get RL-ified without log-probs โ DSRL (noise-latent RL), DiWA โ the seed for ICLR 2026's explosion of RL-for-flow-matching.
- Contact / force becomes first-class โ TA-VLA (single torque token in decoder), DexSkin, Tactile Beyond Pixels, UniFP (Best Paper).
- World models move inside training โ DreamGen trains policies inside video-world-model rollouts; NVIDIA's GR00T N1.6 + Newton + Cosmos are the ecosystem analog.
flowchart LR
V[Vision] --> B[VLM Backbone]
L[Task instruction] --> B
B --> H[High-Level<br/>Subtask Predictor]
B --> AE[Flow-Matching<br/>Action Expert]
H --> AE
AE --> A[Continuous actions]
WD[Web data<br/>+ detection<br/>+ subtask text] -. co-training .-> B
ฯ0.5 (Physical Intelligence, CoRL 2025 Oral + Best Paper Award Finalist) is the flagship CoRL 2025 paper for this wiki's purposes. It adds a high-level subtask predictor on top of ฯ0's flow-matching action expert, co-trains on heterogeneous data (multi-robot teleop + web VQA + detection + subtask semantic prediction), and is the first VLA to demonstrate long-horizon manipulation (cleaning kitchens, bathrooms) in entirely unseen homes.
Every subsequent ฯ release builds on ฯ0.5: ฯ0.6 (Nov 2025, Gemma3-4B + Knowledge Insulation), ฯ*0.6 + RECAP (Nov 2025, RL-from-experience), ฯ0.7 (Apr 2026, MEM history + subgoal-image world model + metadata prompting). See ฯ series evolution.
- VLA Architecture
- Training & Inference Recipes
- Dexterous / Humanoid / Whole-Body Manipulation
- Cross-Embodiment from Human Video
- Diffusion / Flow Policies + RL
- World Models for Policy Learning
- Data & Benchmarks
- Trends & Key Findings
- CoRL 2025 โ ICLR 2026 Lineage
- Reading List
| Paper | Novelty |
|---|---|
| ฯ0.5 (Oral + Best Paper Finalist) | Hierarchical VLA = language subtask head + flow-matching action expert; co-training on heterogeneous sources unlocks open-world generalization |
| DexVLA | ~1B-parameter plug-in diffusion action expert attached to a VLM backbone; cross-embodiment training across arms, dex hands, bimanual |
| TA-VLA | Single torque-history token injected in the decoder raises contact-rich success from โค25% to 85โ90% with no regression on non-contact tasks |
| Streaming Flow Policy (Oral) | Treats the action chunk as a single point on a longer flow trajectory โ can be streamed and extended for low-latency rollout |
Also worth knowing: FLOWER (efficient VLA for consumer hardware), 3DS-VLA (3D spatial features inside a VLA), Long-VLA (long-horizon planning tokens), OneTwoVLA (reasoning + acting dual mode).
| Paper | Novelty |
|---|---|
| ECoT-Lite | Ablation-driven lightweight recipe for embodied chain-of-thought: isolates which sub-CoT steps (plan / subtask / gripper pose / bbox) actually drive downstream gains |
| RoboMonkey | Test-time best-of-N: sample many VLA proposals, score with a learned verifier, execute the best โ brings LLM-style inference-scaling to VLAs |
| DemoSpeedup (Tsinghua) | Entropy-guided demonstration re-timing: preserves low-entropy precise segments, accelerates high-entropy ones โ 1.7โ3ร policy speed-up with no accuracy loss |
| SAIL | Streaming action inference โ extends previous trajectory instead of regenerating, cutting latency and drift |
| RICL | First convincing rollout-based in-context learning for pretrained VLAs (no fine-tune) |
| ControlVLA | Few-shot object-centric adaptation via lightweight control adapters |
| Paper | Novelty |
|---|---|
| DexUMI (Finalist) | Human hand as universal manipulation interface โ wearable sensorized glove directly supervises robot hands |
| ClutterDexGrasp (Oral) | First zero-shot sim-to-real closed-loop target-oriented dexterous grasping in clutter |
| Visual Imitation โ Humanoid (Best Student Paper) | Internet video โ simulated humanoid โ real humanoid; contextual whole-body policies distilled from everyday footage |
| UniFP (Best Paper) | Single policy outputting both position and force targets for legged loco-manipulation |
| Fabrica (Best Paper) | Dual-arm assembly of general multi-part objects combining planning + learned manipulation |
| DexSkin (Oral) | Full-coverage conformable robotic skin + contact-rich dex LfD |
| KineSoft (Oral) | Proprioception-only imitation on soft hands |
| BEHAVIOR Robot Suite (Stanford) | Open-source whole-body bimanual mobile platform + household task benchmark |
| Mobi-ฯ | Mobilizes fixed-base manipulation policies via a learned whole-body controller |
| GraspVLA (GalBot) | Grasping foundation model pretrained on billion-scale synthetic action data |
Human video is the default data source at CoRL 2025 for scaling across embodiments:
- X-Sim (Oral) โ real-to-sim-to-real via object motion as the embodiment-agnostic supervisory signal (Cornell).
- UniSkill โ embodiment-invariant skill embedding learned from human video.
- ImMimic (Oral) โ interpolation between human and robot spaces for cross-domain imitation.
- DexUMI โ already listed; wearable human-hand interface is a human-video analog for dexterous data.
The 2025 insight: diffusion policies have latent structure you can do RL on without ever computing action log-probs.
flowchart LR
N[Initial noise z0<br/>RL action space] --> D[Frozen diffusion policy]
D --> A[Action chunk]
A --> R[Real-world reward]
R -- update z0 policy --> N
- DSRL (Oral, Finalist) โ the headline result: initial-noise latent encodes trajectory info; RL in that latent space steers a frozen diffusion policy cheaply and stably.
- DiWA โ adapts a pretrained diffusion policy using rollouts in a learned world model.
- Belief-Conditioned One-Step Diffusion (Oral) โ trades denoising steps for belief conditioning to hit real-time budgets.
- LatentToM (Oral) โ decentralized diffusion policy architecture for cooperative multi-robot.
- D-CODA โ diffusion-based data augmentation targeting bimanual coordination scarcity.
This whole direction is the seed of ICLR 2026's explosion of RL-for-flow-matching: RECAP, RL Tokens, SimpleVLA-RL, VLA-RFT, Stage-Aware RL.
| Paper | Novelty |
|---|---|
| DreamGen (NVIDIA) | Trains robot policies inside a video-world-model's rollouts at scale |
| LaDi-WM | Latent-diffusion world model tailored to manipulation prediction |
| FLARE | Implicit (non-reconstructive) world-model auxiliary objective |
| ParticleFormer (Stanford) | 3D particle-based dynamics handling cloth / granular / rigid |
Ecosystem signal: NVIDIA launched Isaac GR00T N1.6 + Newton physics engine + Cosmos simulation/generation tools at CoRL 2025. These thread into ICLR 2026 via Ctrl-World, Cosmos Policy, WorldGym.
- ManipBench โ 12,617 MCQs across 33 VLMs / 10 families, first benchmark squarely on low-level manipulation reasoning (objectโobject interaction, deformables). Foreshadows ICLR 2026's VLM4VLA finding that VLM scores โ manipulation skill.
- BEHAVIOR Robot Suite โ reference platform + benchmark for whole-body household tasks.
- Tactile Beyond Pixels (Meta FAIR + GT) โ pretrains multimodal touch representations from ~1M contact-rich interactions across four tactile modalities.
- Hierarchy wins in the open world โ ฯ0.5's subtask-head + action-expert split is the template every serious generalist VLA now follows.
- Test-time scaling arrives for VLAs โ RoboMonkey-style best-of-N is cheap and composable.
- Human video = cross-embodiment data โ wearable interfaces, retargeting, object-motion supervision all converge.
- RL for diffusion policies, sans log-probs โ DSRL is the seed idea; every ICLR 2026 RL-for-VLA paper can trace an ancestor here.
- Contact/force is now first-class โ single-token torque conditioning, full-hand tactile, unified force+position.
- World models inside training โ DreamGen and NVIDIA's stack are the production analog.
Direct descendant threads (links go to ICLR 2026 pages):
| CoRL 2025 seed | ICLR 2026 offspring |
|---|---|
| ฯ0.5 (Oral) | ฯ0.6 โ ฯ*0.6 + RECAP โ ฯ0.7. See ฯ series evolution. |
| DSRL (noise-latent RL) | RECAP, RL Tokens, SimpleVLA-RL, VLA-RFT, Stage-Aware RL, VITA |
| DreamGen / LaDi-WM / FLARE | Ctrl-World, Cosmos Policy, WorldGym |
| ECoT-Lite | Embodied-R1, Actions as Language, InstructVLA |
| DexVLA (plug-in diffusion expert) | Directly challenged by Discrete Diffusion VLA, Unified Diffusion VLA, dVLA |
| UniSkill / X-Sim / ImMimic | X-VLA, XR-1, UniVLA |
| DexUMI / DexSkin / ClutterDexGrasp | DexNDM, UniHM, RFS |
| BEHAVIOR Robot Suite / Mobi-ฯ | WholeBodyVLA, RoboCasa365, RoboArena |
| ManipBench | VLM4VLA (the "VLM score โ manipulation skill" finding) |
Papers I could not confirm as CoRL 2025 (intentionally absent): Octo (ICRA 2024), OpenVLA (CoRL 2024), RDT-1B (ICLR 2025), HPT (NeurIPS 2024). Helix (Figure AI) and Gemini Robotics (DeepMind) were industry demos around the venue, not accepted CoRL 2025 papers.
| Priority | Papers |
|---|---|
| Must read | ฯ0.5, DSRL, DreamGen, DexUMI |
| Award winners | Fabrica, UniFP, Visual Imitation โ Humanoid |
| Architecture | DexVLA, Streaming Flow Policy, TA-VLA |
| Training / inference | ECoT-Lite, RoboMonkey |
| Cross-embodiment | X-Sim |
| Dexterous | ClutterDexGrasp, DexUMI |
| Benchmarks | ManipBench |
- CoRL 2025 official site
- CoRL 2025 awards
- Stanford AI Lab CoRL 2025 summary
- Curated 221-paper aggregator
- Paper Copilot โ CoRL 2025 stats
- Paper Copilot โ CoRL 2025 paper list
- Lumeny attendee trends summary
- NVIDIA at CoRL 2025 (GR00T N1.6 + Newton + Cosmos)
- The Robot Report โ Newton / GR00T N1.6 launch
Per-paper links appear on each paper's page.