ICML 2026 Scaling by Diversified Experience for - Heungwoo/research GitHub Wiki
Scaling by Diversified Experience for Vision-Language-Action Models — SyVLA: intention decoupling + similar-sample-guided RL
Venue: ICML 2026 (Poster) Category: RL for VLA
Problem
Vision-Language-Action (VLA) models face significant challenges in real-world deployment due to two issues: the entanglement of high-level reasoning with low-level control, and the instability of policy optimization. When control-relevant features are not cleanly separated from reasoning context, and when reinforcement learning updates are unstable under distribution shift, deployed policies generalize poorly and can erode the model's underlying vision-language capabilities.
Method
The authors introduce SyVLA, a robust VLA model trained with diversified experiences. It combines two components:
- An Intention Decoupling algorithm to isolate control-relevant features from reasoning contexts, disentangling high-level reasoning from low-level action prediction.
- A similar-sample guided RL pipeline to stabilize policy updates and mitigate distribution shift during training.
flowchart LR
A[Vision + Language input] --> B[Intention Decoupling]
B --> C[Reasoning context]
B --> D[Control-relevant features]
D --> E[Similar-sample guided RL]
E --> F[Stabilized policy update]
C --> F
F --> G[SyVLA policy: actions]
(Schematic derived from the abstract; the paper's own figures were not available.)
Results
Extensive experiments on real-world robotic tasks and multi-modal benchmarks demonstrate that SyVLA achieves superior task success rates and stronger out-of-distribution generalization compared to existing methods, while effectively preserving core vision-language capabilities. (Specific numeric results were not available from the public ICML abstract.)
Significance
SyVLA targets two recurring failure modes of VLA training at once — reasoning/control entanglement and RL instability — rather than treating them separately. By decoupling intention from reasoning and grounding policy updates in similar samples, it aims for robust OOD generalization without sacrificing the pretrained vision-language competence that makes generalist VLAs valuable.
Links
- ICML 2026: https://icml.cc/virtual/2026/poster/62749
← Back to ICML-2026