ICLR 2026 PLD - Heungwoo/research GitHub Wiki

PLD โ€” Probe, Learn, Distill (Self-Improving VLA via Residual RL)

Venue: ICLR 2026 ยท Project: wenlixiao.com/self-improve-VLA-PLD Category: RL for VLA Trend tag: Trend 3

Approach diagram

flowchart LR
  Step1[1. PROBE<br/>Distribution-aware data collection<br/>find failure cases of base VLA] --> Step2[2. LEARN<br/>Train lightweight residual actors<br/>off-policy RL SAC on probed failures<br/>BACKBONE FROZEN]
  Step2 --> Step3[3. DISTILL<br/>Distill residual-corrected behavior<br/>back into base VLA]
  Step3 --> Out[Self-improved VLA<br/>~99% LIBERO ยท 100% real Franka & YAM]
  Step3 -. iterate .-> Step1
Loading

Problem

Standard RL fine-tuning of a pretrained VLA destabilizes the backbone โ€” on-policy gradients pull the full VLA policy away from its pretrained distribution, causing catastrophic capability loss. Freezing the backbone and only training an action head doesn't help either โ€” the information bottleneck is too tight. (Base policies evaluated are ฯ€0 (flow-matching action head) and OpenVLA (autoregressive action tokens); the paper does not quote a single fixed parameter count.)

Method

Three stages:

  1. Probe โ€” identify failure cases via distribution-aware data collection
  2. Learn โ€” train lightweight residual actors on top of the frozen base VLA using sample-efficient off-policy RL (SAC, Gaussian policy) on probed failures; the residual outputs zero most of the time and only intervenes in failure states (only the residual updates online โ†’ backbone stable)
  3. Distill โ€” distill the residual-corrected behavior back into the base VLA โ†’ self-improved model

Loop is repeatable.

Results

  • ~99% task success on LIBERO (near saturation)
  • >50% performance gain on SimplerEnv
  • 100% success on real-world Franka and YAM dexterous manipulation tasks

Among the strongest VLA RL numbers published.

Significance

Most compelling single demonstration that RL fine-tuning of VLAs is practical. Probe-Learn-Distill is repeatable and generalizes across platforms. LIBERO's saturation after PLD is the main trigger for the benchmark-generation work of RoboArena โˆž and WorldGym.

Links

Related pages

โ† Back to ICLR-2026 ยท Topic: RL

โš ๏ธ **GitHub.com Fallback** โš ๏ธ