CoRL 2025 VLA Manipulation Survey - Heungwoo/research GitHub Wiki

CoRL 2025 โ€” VLA & Manipulation Survey

Compiled April 2026 (in retrospect). Focus: VLA architecture, training & inference recipes, dexterous / humanoid manipulation, cross-embodiment, diffusion & flow policies, world models, benchmarks.

TL;DR

CoRL 2025 (Seoul, Sept 27โ€“30) accepted 221 papers out of 263 submissions (84.03% acceptance), of which 42 were orals. From the VLA/manipulation perspective, six trends stand out:

  1. Hierarchical VLAs beat monolithic ones for open-world generalization โ€” exemplified by ฯ€0.5 (CoRL 2025 Oral + Best Paper Award Finalist), which sets up the entire ฯ€ series that the field now benchmarks against.
  2. Test-time scaling + trajectory streaming โ€” RoboMonkey (best-of-N with verifier), Streaming Flow Policy, DemoSpeedup, SAIL.
  3. Human video as the default cross-embodiment data source โ€” DexUMI, Visual Imitation โ†’ Humanoid, UniSkill, ImMimic, X-Sim.
  4. Diffusion / flow policies get RL-ified without log-probs โ€” DSRL (noise-latent RL), DiWA โ€” the seed for ICLR 2026's explosion of RL-for-flow-matching.
  5. Contact / force becomes first-class โ€” TA-VLA (single torque token in decoder), DexSkin, Tactile Beyond Pixels, UniFP (Best Paper).
  6. World models move inside training โ€” DreamGen trains policies inside video-world-model rollouts; NVIDIA's GR00T N1.6 + Newton + Cosmos are the ecosystem analog.

Flagship baseline: ฯ€0.5

flowchart LR
  V[Vision] --> B[VLM Backbone]
  L[Task instruction] --> B
  B --> H[High-Level<br/>Subtask Predictor]
  B --> AE[Flow-Matching<br/>Action Expert]
  H --> AE
  AE --> A[Continuous actions]
  WD[Web data<br/>+ detection<br/>+ subtask text] -. co-training .-> B
Loading

ฯ€0.5 (Physical Intelligence, CoRL 2025 Oral + Best Paper Award Finalist) is the flagship CoRL 2025 paper for this wiki's purposes. It adds a high-level subtask predictor on top of ฯ€0's flow-matching action expert, co-trains on heterogeneous data (multi-robot teleop + web VQA + detection + subtask semantic prediction), and is the first VLA to demonstrate long-horizon manipulation (cleaning kitchens, bathrooms) in entirely unseen homes.

Every subsequent ฯ€ release builds on ฯ€0.5: ฯ€0.6 (Nov 2025, Gemma3-4B + Knowledge Insulation), ฯ€*0.6 + RECAP (Nov 2025, RL-from-experience), ฯ€0.7 (Apr 2026, MEM history + subgoal-image world model + metadata prompting). See ฯ€ series evolution.

Section index

  1. VLA Architecture
  2. Training & Inference Recipes
  3. Dexterous / Humanoid / Whole-Body Manipulation
  4. Cross-Embodiment from Human Video
  5. Diffusion / Flow Policies + RL
  6. World Models for Policy Learning
  7. Data & Benchmarks
  8. Trends & Key Findings
  9. CoRL 2025 โ†’ ICLR 2026 Lineage
  10. Reading List

1. VLA Architecture

Paper Novelty
ฯ€0.5 (Oral + Best Paper Finalist) Hierarchical VLA = language subtask head + flow-matching action expert; co-training on heterogeneous sources unlocks open-world generalization
DexVLA ~1B-parameter plug-in diffusion action expert attached to a VLM backbone; cross-embodiment training across arms, dex hands, bimanual
TA-VLA Single torque-history token injected in the decoder raises contact-rich success from โ‰ค25% to 85โ€“90% with no regression on non-contact tasks
Streaming Flow Policy (Oral) Treats the action chunk as a single point on a longer flow trajectory โ€” can be streamed and extended for low-latency rollout

Also worth knowing: FLOWER (efficient VLA for consumer hardware), 3DS-VLA (3D spatial features inside a VLA), Long-VLA (long-horizon planning tokens), OneTwoVLA (reasoning + acting dual mode).


2. Training & Inference Recipes

Paper Novelty
ECoT-Lite Ablation-driven lightweight recipe for embodied chain-of-thought: isolates which sub-CoT steps (plan / subtask / gripper pose / bbox) actually drive downstream gains
RoboMonkey Test-time best-of-N: sample many VLA proposals, score with a learned verifier, execute the best โ€” brings LLM-style inference-scaling to VLAs
DemoSpeedup (Tsinghua) Entropy-guided demonstration re-timing: preserves low-entropy precise segments, accelerates high-entropy ones โ†’ 1.7โ€“3ร— policy speed-up with no accuracy loss
SAIL Streaming action inference โ€” extends previous trajectory instead of regenerating, cutting latency and drift
RICL First convincing rollout-based in-context learning for pretrained VLAs (no fine-tune)
ControlVLA Few-shot object-centric adaptation via lightweight control adapters

3. Dexterous / Humanoid / Whole-Body Manipulation

Paper Novelty
DexUMI (Finalist) Human hand as universal manipulation interface โ€” wearable sensorized glove directly supervises robot hands
ClutterDexGrasp (Oral) First zero-shot sim-to-real closed-loop target-oriented dexterous grasping in clutter
Visual Imitation โ†’ Humanoid (Best Student Paper) Internet video โ†’ simulated humanoid โ†’ real humanoid; contextual whole-body policies distilled from everyday footage
UniFP (Best Paper) Single policy outputting both position and force targets for legged loco-manipulation
Fabrica (Best Paper) Dual-arm assembly of general multi-part objects combining planning + learned manipulation
DexSkin (Oral) Full-coverage conformable robotic skin + contact-rich dex LfD
KineSoft (Oral) Proprioception-only imitation on soft hands
BEHAVIOR Robot Suite (Stanford) Open-source whole-body bimanual mobile platform + household task benchmark
Mobi-ฯ€ Mobilizes fixed-base manipulation policies via a learned whole-body controller
GraspVLA (GalBot) Grasping foundation model pretrained on billion-scale synthetic action data

4. Cross-Embodiment from Human Video

Human video is the default data source at CoRL 2025 for scaling across embodiments:

  • X-Sim (Oral) โ€” real-to-sim-to-real via object motion as the embodiment-agnostic supervisory signal (Cornell).
  • UniSkill โ€” embodiment-invariant skill embedding learned from human video.
  • ImMimic (Oral) โ€” interpolation between human and robot spaces for cross-domain imitation.
  • DexUMI โ€” already listed; wearable human-hand interface is a human-video analog for dexterous data.

5. Diffusion / Flow Policies + RL

The 2025 insight: diffusion policies have latent structure you can do RL on without ever computing action log-probs.

flowchart LR
  N[Initial noise z0<br/>RL action space] --> D[Frozen diffusion policy]
  D --> A[Action chunk]
  A --> R[Real-world reward]
  R -- update z0 policy --> N
Loading
  • DSRL (Oral, Finalist) โ€” the headline result: initial-noise latent encodes trajectory info; RL in that latent space steers a frozen diffusion policy cheaply and stably.
  • DiWA โ€” adapts a pretrained diffusion policy using rollouts in a learned world model.
  • Belief-Conditioned One-Step Diffusion (Oral) โ€” trades denoising steps for belief conditioning to hit real-time budgets.
  • LatentToM (Oral) โ€” decentralized diffusion policy architecture for cooperative multi-robot.
  • D-CODA โ€” diffusion-based data augmentation targeting bimanual coordination scarcity.

This whole direction is the seed of ICLR 2026's explosion of RL-for-flow-matching: RECAP, RL Tokens, SimpleVLA-RL, VLA-RFT, Stage-Aware RL.


6. World Models for Policy Learning

Paper Novelty
DreamGen (NVIDIA) Trains robot policies inside a video-world-model's rollouts at scale
LaDi-WM Latent-diffusion world model tailored to manipulation prediction
FLARE Implicit (non-reconstructive) world-model auxiliary objective
ParticleFormer (Stanford) 3D particle-based dynamics handling cloth / granular / rigid

Ecosystem signal: NVIDIA launched Isaac GR00T N1.6 + Newton physics engine + Cosmos simulation/generation tools at CoRL 2025. These thread into ICLR 2026 via Ctrl-World, Cosmos Policy, WorldGym.


7. Data & Benchmarks

  • ManipBench โ€” 12,617 MCQs across 33 VLMs / 10 families, first benchmark squarely on low-level manipulation reasoning (objectโ€“object interaction, deformables). Foreshadows ICLR 2026's VLM4VLA finding that VLM scores โ‰  manipulation skill.
  • BEHAVIOR Robot Suite โ€” reference platform + benchmark for whole-body household tasks.
  • Tactile Beyond Pixels (Meta FAIR + GT) โ€” pretrains multimodal touch representations from ~1M contact-rich interactions across four tactile modalities.

8. Trends & Key Findings

  1. Hierarchy wins in the open world โ€” ฯ€0.5's subtask-head + action-expert split is the template every serious generalist VLA now follows.
  2. Test-time scaling arrives for VLAs โ€” RoboMonkey-style best-of-N is cheap and composable.
  3. Human video = cross-embodiment data โ€” wearable interfaces, retargeting, object-motion supervision all converge.
  4. RL for diffusion policies, sans log-probs โ€” DSRL is the seed idea; every ICLR 2026 RL-for-VLA paper can trace an ancestor here.
  5. Contact/force is now first-class โ€” single-token torque conditioning, full-hand tactile, unified force+position.
  6. World models inside training โ€” DreamGen and NVIDIA's stack are the production analog.

9. CoRL 2025 โ†’ ICLR 2026 Lineage

Direct descendant threads (links go to ICLR 2026 pages):

CoRL 2025 seed ICLR 2026 offspring
ฯ€0.5 (Oral) ฯ€0.6 โ†’ ฯ€*0.6 + RECAP โ†’ ฯ€0.7. See ฯ€ series evolution.
DSRL (noise-latent RL) RECAP, RL Tokens, SimpleVLA-RL, VLA-RFT, Stage-Aware RL, VITA
DreamGen / LaDi-WM / FLARE Ctrl-World, Cosmos Policy, WorldGym
ECoT-Lite Embodied-R1, Actions as Language, InstructVLA
DexVLA (plug-in diffusion expert) Directly challenged by Discrete Diffusion VLA, Unified Diffusion VLA, dVLA
UniSkill / X-Sim / ImMimic X-VLA, XR-1, UniVLA
DexUMI / DexSkin / ClutterDexGrasp DexNDM, UniHM, RFS
BEHAVIOR Robot Suite / Mobi-ฯ€ WholeBodyVLA, RoboCasa365, RoboArena
ManipBench VLM4VLA (the "VLM score โ‰  manipulation skill" finding)

Papers I could not confirm as CoRL 2025 (intentionally absent): Octo (ICRA 2024), OpenVLA (CoRL 2024), RDT-1B (ICLR 2025), HPT (NeurIPS 2024). Helix (Figure AI) and Gemini Robotics (DeepMind) were industry demos around the venue, not accepted CoRL 2025 papers.


10. Reading List

Priority Papers
Must read ฯ€0.5, DSRL, DreamGen, DexUMI
Award winners Fabrica, UniFP, Visual Imitation โ†’ Humanoid
Architecture DexVLA, Streaming Flow Policy, TA-VLA
Training / inference ECoT-Lite, RoboMonkey
Cross-embodiment X-Sim
Dexterous ClutterDexGrasp, DexUMI
Benchmarks ManipBench

Sources

Per-paper links appear on each paper's page.

โ† Back to CoRL-2025 ยท CoRL ยท Home

โš ๏ธ **GitHub.com Fallback** โš ๏ธ