CoRL 2025 - Heungwoo/research GitHub Wiki
CoRL 2025
Conference on Robot Learning 2025 โ Seoul, Sept 27โ30, 2025.
263 accepted papers (42 orals + 221 posters). VLA/manipulation was the dominant thread; NVIDIA launched Isaac GR00T N1.6 + the Newton physics engine at the venue.
Surveys hosted in this wiki
- VLA & Manipulation Survey (CoRL 2025) โ categorized summary of CoRL 2025's VLA/manipulation papers plus trends and a CoRL 2025 โ ICLR 2026 lineage map.
Awards
| Award | Paper |
|---|---|
| Best Paper | [Fabrica](/Heungwoo/research/wiki/CoRL-2025-Fabrica) โ dual-arm general multi-part assembly |
| Best Paper | [UniFP](/Heungwoo/research/wiki/CoRL-2025-UniFP) โ unified position+force policy for legged loco-manipulation |
| Best Student Paper | [Visual Imitation โ Humanoid](/Heungwoo/research/wiki/CoRL-2025-Visual-Imitation-Humanoid) โ contextual humanoid control from internet video |
| Best Paper Finalist | [DexUMI](/Heungwoo/research/wiki/CoRL-2025-DexUMI), [DSRL](/Heungwoo/research/wiki/CoRL-2025-DSRL), LocoFormer, The Sound of Simulation, [ฯ0.5](/Heungwoo/research/wiki/CoRL-2025-pi05) |
Quick paper index (by category)
Flagship baseline
- ฯ0.5 (Oral) โ Physical Intelligence's hierarchical VLA with co-training; sets up the entire 2025โ2026 ฯ series. See ฯ series evolution.
VLA architecture
- DexVLA โ plug-in ~1B diffusion action expert atop a VLM
- TA-VLA โ single torque-history token in the decoder for contact-rich tasks
- Streaming Flow Policy (Oral) โ action chunk = point on a longer flow
Training & inference recipes
- ECoT-Lite โ which parts of embodied chain-of-thought actually matter
- RoboMonkey โ best-of-N VLA sampling with a learned verifier
Dexterous / humanoid / whole-body
- DexUMI (Finalist) โ wearable "universal manipulation interface"
- ClutterDexGrasp (Oral) โ zero-shot sim-to-real closed-loop dex grasping in clutter
- UniFP (Best Paper) โ unified position+force legged loco-manipulation
- Visual Imitation โ Humanoid (Best Student Paper) โ everyday video โ humanoid skills
- Fabrica (Best Paper) โ dual-arm multi-part assembly
Cross-embodiment from human video
- X-Sim (Oral) โ real-to-sim-to-real via object motion
Diffusion / flow policies + RL
- DSRL (Oral, Finalist) โ RL in the initial-noise latent space of a frozen diffusion policy
World models for policy
- DreamGen โ policy training inside a video world model (NVIDIA)
Data & benchmarks
- ManipBench โ first VLM benchmark targeting low-level manipulation reasoning
Trends (6)
- Hierarchical VLAs beat monolithic VLAs for open-world generalization (ฯ0.5, OneTwoVLA, Long-VLA).
- Test-time scaling + trajectory streaming are the cheap latency/quality levers (RoboMonkey, Streaming Flow Policy, DemoSpeedup, SAIL).
- Human video is the default cross-embodiment data source (DexUMI, Visual Imitation โ Humanoid, UniSkill, ImMimic, X-Sim).
- Diffusion / flow policies get RL-ified without log-probs (DSRL, DiWA) โ prefigures ICLR 2026's RECAP, RL Tokens, VLA-RFT, SimpleVLA-RL.
- Contact / force is finally modeled inside VLAs (TA-VLA, DexSkin, UniFP, KineSoft, Tactile Beyond Pixels).
- World models move inside (not adjacent to) training pipelines (DreamGen, LaDi-WM, FLARE, ParticleFormer) โ with NVIDIA GR00T N1.6 + Newton as the ecosystem signal.