CoRL 2025 - Heungwoo/research GitHub Wiki

CoRL 2025

Conference on Robot Learning 2025 โ€” Seoul, Sept 27โ€“30, 2025.

263 accepted papers (42 orals + 221 posters). VLA/manipulation was the dominant thread; NVIDIA launched Isaac GR00T N1.6 + the Newton physics engine at the venue.

Surveys hosted in this wiki

Awards

Award Paper
Best Paper [Fabrica](/Heungwoo/research/wiki/CoRL-2025-Fabrica) โ€” dual-arm general multi-part assembly
Best Paper [UniFP](/Heungwoo/research/wiki/CoRL-2025-UniFP) โ€” unified position+force policy for legged loco-manipulation
Best Student Paper [Visual Imitation โ†’ Humanoid](/Heungwoo/research/wiki/CoRL-2025-Visual-Imitation-Humanoid) โ€” contextual humanoid control from internet video
Best Paper Finalist [DexUMI](/Heungwoo/research/wiki/CoRL-2025-DexUMI), [DSRL](/Heungwoo/research/wiki/CoRL-2025-DSRL), LocoFormer, The Sound of Simulation, [ฯ€0.5](/Heungwoo/research/wiki/CoRL-2025-pi05)

Quick paper index (by category)

Flagship baseline

  • ฯ€0.5 (Oral) โ€” Physical Intelligence's hierarchical VLA with co-training; sets up the entire 2025โ†’2026 ฯ€ series. See ฯ€ series evolution.

VLA architecture

  • DexVLA โ€” plug-in ~1B diffusion action expert atop a VLM
  • TA-VLA โ€” single torque-history token in the decoder for contact-rich tasks
  • Streaming Flow Policy (Oral) โ€” action chunk = point on a longer flow

Training & inference recipes

  • ECoT-Lite โ€” which parts of embodied chain-of-thought actually matter
  • RoboMonkey โ€” best-of-N VLA sampling with a learned verifier

Dexterous / humanoid / whole-body

  • DexUMI (Finalist) โ€” wearable "universal manipulation interface"
  • ClutterDexGrasp (Oral) โ€” zero-shot sim-to-real closed-loop dex grasping in clutter
  • UniFP (Best Paper) โ€” unified position+force legged loco-manipulation
  • Visual Imitation โ†’ Humanoid (Best Student Paper) โ€” everyday video โ†’ humanoid skills
  • Fabrica (Best Paper) โ€” dual-arm multi-part assembly

Cross-embodiment from human video

  • X-Sim (Oral) โ€” real-to-sim-to-real via object motion

Diffusion / flow policies + RL

  • DSRL (Oral, Finalist) โ€” RL in the initial-noise latent space of a frozen diffusion policy

World models for policy

  • DreamGen โ€” policy training inside a video world model (NVIDIA)

Data & benchmarks

  • ManipBench โ€” first VLM benchmark targeting low-level manipulation reasoning

Trends (6)

  1. Hierarchical VLAs beat monolithic VLAs for open-world generalization (ฯ€0.5, OneTwoVLA, Long-VLA).
  2. Test-time scaling + trajectory streaming are the cheap latency/quality levers (RoboMonkey, Streaming Flow Policy, DemoSpeedup, SAIL).
  3. Human video is the default cross-embodiment data source (DexUMI, Visual Imitation โ†’ Humanoid, UniSkill, ImMimic, X-Sim).
  4. Diffusion / flow policies get RL-ified without log-probs (DSRL, DiWA) โ€” prefigures ICLR 2026's RECAP, RL Tokens, VLA-RFT, SimpleVLA-RL.
  5. Contact / force is finally modeled inside VLAs (TA-VLA, DexSkin, UniFP, KineSoft, Tactile Beyond Pixels).
  6. World models move inside (not adjacent to) training pipelines (DreamGen, LaDi-WM, FLARE, ParticleFormer) โ€” with NVIDIA GR00T N1.6 + Newton as the ecosystem signal.

โ† Back to CoRL ยท Home