ICLR 2026 Stage Aware RL - Heungwoo/research GitHub Wiki

STARE-VLA: Progressive Stage-Aware Reinforcement for VLA Fine-Tuning

Venue: ICLR 2026 Category: RL for VLA Trend tag: Trend 3 arXiv: 2512.05107 (Dec 2025, v2) Authors: Feng Xu, Guangyao Zhai, Xin Kong, Tingzhong Fu, Daniel F. N. Gordon, Xueli An, Benjamin Busam (Huawei Munich Research Center; Imperial College London; TUM)

Approach diagram

flowchart LR
  T[Manipulation task] --> S1[Stage 1: REACH]
  S1 --> S2[Stage 2: GRASP]
  S2 --> S3[Stage 3: TRANSPORT]
  S3 --> S4[Stage 4: PLACE]
  S1 -.dense reward.-> R[Stage-aligned reward signal]
  S2 -.dense reward.-> R
  S3 -.dense reward.-> R
  S4 -.dense reward.-> R
  R --> RL[STA-TPO offline preference<br/>+ STA-PPO online interaction]
Loading

Reach→Grasp→Transport→Place is the canonical pick-and-place decomposition; other tasks use stages such as Push/Pull, Lift, Upright, and Goal. The number of stages varies (typically 3–5) per task.

Problem

Existing RL fine-tuning treats long-horizon action trajectories as linguistic sequences and applies trajectory-level objectives — Trajectory-wise Preference Optimization (TPO) or PPO — over the whole trajectory. This gives coarse credit assignment and unstable training. The paper's key observation: unlike language, where meaning is preserved despite flexible word order, action trajectories progress through causally chained stages with different learning difficulties, motivating progressive stage-wise optimization.

Method

STARE (Stage-Aware Reinforcement) is a module that decomposes a long-horizon trajectory into semantically meaningful stages and emits dense, interpretable, stage-aligned reward signals. Segmentation is rule-based / event-driven: stage boundaries are set by geometric and semantic thresholds (δ_k) on end-effector translation/orientation and object state (e.g., Reach→Grasp when the gripper contacts the object; Transport→Place when the object is within a distance margin of the goal) rather than arbitrary temporal cuts.

STARE is plugged into both optimizers:

  • STA-TPO — offline, stage-wise preference learning.
  • STA-PPO — online, intra-stage interaction.

These are composed in the IPI pipeline (Imitation → Preference → Interaction): SFT initialization, then STA-TPO offline preference, then STA-PPO online interaction. Base VLAs fine-tuned: OpenVLA-7B and π0.5_base.

Results

Evaluated on SimplerEnv and ManiSkill3 (OpenVLA-7B backbone). The full IPI pipeline reaches state-of-the-art 98.0% on SimplerEnv and 96.4% on ManiSkill3.

Method SimplerEnv ManiSkill3
SFT 41.7% 15.0%
GRAPE 43.9% 18.0%
STA-TPO 51.6% 20.8%
RL4VLA 92.5% 70.5%
STA-PPO 94.6% 93.4%
IPI (full) 98.0% 96.4%

Ablation: disabling STARE at the final high-precision stages (Place, Upright) causes >20% success drops, showing stage-aligned signal matters most where precision is critical.

Significance

Reframes VLA RL fine-tuning from trajectory-level to stage-level credit assignment, addressing the instability of treating actions as flat token sequences. The event-driven stage segmentation needs no hand-tuned dense reward per task, and the IPI ordering (preference before interaction) yields large gains over both pure offline (STA-TPO) and pure online (STA-PPO) RL alone.

Links

Related pages

  • PLD (residual-RL alternative)
  • VITA (reward modeling alternative)

← Back to ICLR-2026 · Topic: RL

⚠️ **GitHub.com Fallback** ⚠️