ICLR 2026 PLD - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 ยท Project: wenlixiao.com/self-improve-VLA-PLD Category: RL for VLA Trend tag: Trend 3
flowchart LR
Step1[1. PROBE<br/>Distribution-aware data collection<br/>find failure cases of base VLA] --> Step2[2. LEARN<br/>Train lightweight residual actors<br/>off-policy RL SAC on probed failures<br/>BACKBONE FROZEN]
Step2 --> Step3[3. DISTILL<br/>Distill residual-corrected behavior<br/>back into base VLA]
Step3 --> Out[Self-improved VLA<br/>~99% LIBERO ยท 100% real Franka & YAM]
Step3 -. iterate .-> Step1
Standard RL fine-tuning of a pretrained VLA destabilizes the backbone โ on-policy gradients pull the full VLA policy away from its pretrained distribution, causing catastrophic capability loss. Freezing the backbone and only training an action head doesn't help either โ the information bottleneck is too tight. (Base policies evaluated are ฯ0 (flow-matching action head) and OpenVLA (autoregressive action tokens); the paper does not quote a single fixed parameter count.)
Three stages:
- Probe โ identify failure cases via distribution-aware data collection
- Learn โ train lightweight residual actors on top of the frozen base VLA using sample-efficient off-policy RL (SAC, Gaussian policy) on probed failures; the residual outputs zero most of the time and only intervenes in failure states (only the residual updates online โ backbone stable)
- Distill โ distill the residual-corrected behavior back into the base VLA โ self-improved model
Loop is repeatable.
- ~99% task success on LIBERO (near saturation)
- >50% performance gain on SimplerEnv
- 100% success on real-world Franka and YAM dexterous manipulation tasks
Among the strongest VLA RL numbers published.
Most compelling single demonstration that RL fine-tuning of VLAs is practical. Probe-Learn-Distill is repeatable and generalizes across platforms. LIBERO's saturation after PLD is the main trigger for the benchmark-generation work of RoboArena โ and WorldGym.
- Project page: https://wenlixiao.com/self-improve-VLA-PLD
- OpenReview: https://openreview.net/forum?id=xqRD4LLY5M
- VLA-RFT (world-model alternative)
- ฯ*0.6 + RECAP (flow-matching-specific alternative)
- RFS (residual-RL applied to dexterity)