ICML 2026 Scaling by Diversified Experience for - Heungwoo/research GitHub Wiki

Scaling by Diversified Experience for Vision-Language-Action Models — SyVLA: intention decoupling + similar-sample-guided RL

Venue: ICML 2026 (Poster) Category: RL for VLA

Problem

Vision-Language-Action (VLA) models face significant challenges in real-world deployment due to two issues: the entanglement of high-level reasoning with low-level control, and the instability of policy optimization. When control-relevant features are not cleanly separated from reasoning context, and when reinforcement learning updates are unstable under distribution shift, deployed policies generalize poorly and can erode the model's underlying vision-language capabilities.

Method

The authors introduce SyVLA, a robust VLA model trained with diversified experiences. It combines two components:

  • An Intention Decoupling algorithm to isolate control-relevant features from reasoning contexts, disentangling high-level reasoning from low-level action prediction.
  • A similar-sample guided RL pipeline to stabilize policy updates and mitigate distribution shift during training.
flowchart LR
    A[Vision + Language input] --> B[Intention Decoupling]
    B --> C[Reasoning context]
    B --> D[Control-relevant features]
    D --> E[Similar-sample guided RL]
    E --> F[Stabilized policy update]
    C --> F
    F --> G[SyVLA policy: actions]

(Schematic derived from the abstract; the paper's own figures were not available.)

Results

Extensive experiments on real-world robotic tasks and multi-modal benchmarks demonstrate that SyVLA achieves superior task success rates and stronger out-of-distribution generalization compared to existing methods, while effectively preserving core vision-language capabilities. (Specific numeric results were not available from the public ICML abstract.)

Significance

SyVLA targets two recurring failure modes of VLA training at once — reasoning/control entanglement and RL instability — rather than treating them separately. By decoupling intention from reasoning and grounding policy updates in similar samples, it aims for robust OOD generalization without sacrificing the pretrained vision-language competence that makes generalist VLAs valuable.

Links

← Back to ICML-2026