CoRL 2026 LUCID - Heungwoo/research GitHub Wiki
CoRL 2026 โ LUCID: Embodiment-Agnostic Intent Models from Human Videos
Venue: CoRL 2026 (Austin, TX, Nov 9โ12) ยท CMU LeCAR Lab. Paper: arXiv 2606.11628. Representative of: decoupled intent/execution from human video โ learn what should happen from internet video, let each embodiment learn how. Companions: Human Video โ Robot Transfer ยท Dexterous Manipulation ยท CoRL 2026 survey.

1. Problem
Robot-learning pipelines lean on costly robot demonstrations or structured human data that is tied to one embodiment. Unstructured internet human video is abundant but has no direct link to robot actions. LUCID (Harsh Gupta, Guanya Shi, Wenzhen Yuan) asks whether what should happen next โ an embodiment-agnostic intent โ can be learned from such video and then executed by any robot.
2. Method
Two stages, decoupled by an explicit intent interface:
- Intent model (from human video). Adapts CoTracker3 into a point-token transformer that predicts forward in time. It conditions on frozen DINOv3 patch tokens (with depth fused via a residual adapter) and predicts short-horizon object 3D-flow trajectories plus a single palm-pose token (position + rotation) over T future steps. Runs closed-loop: re-predicts intent from current observations each step.
- Sensorimotor policy (per embodiment). Goal-conditioned RL (PPO) in Isaac Lab with massively-parallel sim. A privileged teacher (ฯT) is trained with a curriculum (scalar ฯโ[0,1] tightening gravity, perturbations, success tolerances) and adaptive domain randomization following DextrAH-RGB; a student (ฯS) is distilled with a hybrid PPO + MSE objective. The same intent model drives both a LEAP dexterous hand and a parallel-jaw gripper.
3. Results
- Closed-loop vs open-loop: 73% average success (LUCID closed-loop) vs 28% for an open-loop video-prediction baseline (Veo 3.1).
- Web-scraped tasks (stirring, wiping, binning; 20k clips each) evaluated over 10 trials ร 3 scenarios = 30 per task.
- Self-collected tasks (push-T, cable routing; ~1 hr smartphone video each): 19/30 each.
- Cross-embodiment: dexterous hand and parallel-jaw gripper reach the same aggregate success (19/30 each) on push-T and cable routing.
- Data scaling (binning): 1k clips poor โ 5โ10k localizes containers โ 20k highest, steadily improving.
4. Why it matters
Separating intent (learned once, from free internet video) from execution (learned per robot in sim) sidesteps the embodiment-specific demo bottleneck and lets a single intent model port across hands and grippers. The explicit 3D-flow + palm-pose interface is what makes both the cross-embodiment transfer and the closed-loop re-planning tractable.
Limitations (reviewer): (1) Pipeline brittleness โ many perception modules (SAM, DenseTrack3D, ViPE, WiLoR) stack failure points; end-to-end learning would help. (2) Task-condition gap โ tasks without verifiable endpoints loop indefinitely and manual corpus filtering scales poorly. (3) Lossy interface โ 3D-flow + palm-pose discards finger configuration and fine contact detail.
5. Links
- arXiv 2606.11628
- Survey: CoRL 2026 ยท Related: Human Video โ Robot Transfer ยท Dexterous Manipulation
โ Back to CoRL 2026 survey ยท Home