CoRL 2026 VLS - Heungwoo/research GitHub Wiki
CoRL 2026 — VLS: Steering Pretrained Robot Policies via Vision-Language Models
Venue: CoRL 2026 (Austin, TX, Nov 9–12). Paper: arXiv 2602.03973. Representative of: training-free inference-time policy steering — keep the policy frozen; a VLM-generated reward steers action denoising. Companions: RL for VLA · Real-Time Execution · CoRL 2026 survey.
Authors: Shuo Liu, Ishneet Sukhvinder Singh, Yiqing Xu, Jiafei Duan, Ranjay Krishna.

1. Problem
Pretrained diffusion / flow-matching policies fail when the same task moves near an obstacle, onto a shifted support surface, or into mild clutter. These failures are not missing motor skills — they expose imitation learning's coupling of action generation to training-specific spatial configurations and task specifications. Retraining or fine-tuning is costly and conceptually misaligned: the needed behavior already exists in the policy but cannot be selectively summoned at test time.
2. Method
Vision-Language Steering (VLS) treats adaptation as an inference-time control problem over a frozen generative policy — no parameter updates. Given out-of-distribution observation-language inputs, VLS:
- Grounds the OOD scene into task-relevant 3D keypoints (SAM + DINOv2 features).
- Has a VLM synthesize a stage-aware, trajectory-differentiable reward as PyTorch operations.
- Steers denoising with the reward: gradient-based refinement (MCMC), RBF repulsive forces for trajectory diversity, and gradient-free Feynman–Kac resampling.
- Runs closed-loop, with adaptive guidance strength and Schmitt-trigger stage switching driven by reward feedback.
3. Results
- CALVIN: ~+31% success over the frozen base policy; 94% on movable objects and 87% on articulated parts, beating DynaGuide and ITPS by ~15–25 points.
- LIBERO-PRO: π₀.₅ + VLS reaches 36.81% overall, up to +13% under spatial / semantic perturbations vs. the frozen baseline.
- Real Franka: +19% in-distribution (69% avg); holds up under appearance shift; 40% on a novel-mug substitution where the baseline fails.
4. Why it matters
VLS shows that much of a pretrained policy's OOD "failure" is a retrieval problem, not a capability gap: a VLM-authored reward can steer denoising to recover latent skills without any gradient step on the policy. That reframes robustness as an inference-time, plug-in layer usable on top of any frozen diffusion/flow VLA.
Limitations (reviewer): batch sampling, MCMC steps, and Feynman–Kac resampling add heavy per-step inference overhead in the denoising loop; quality hinges on the VLM's reward synthesis and the accuracy of keypoint grounding.
5. Links
- arXiv 2602.03973
- Survey: CoRL 2026 · Related: RL for VLA · Real-Time Execution
← Back to CoRL 2026 survey · Home