ICLR 2026 Stage Aware RL - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 Category: RL for VLA Trend tag: Trend 3 arXiv: 2512.05107 (Dec 2025, v2) Authors: Feng Xu, Guangyao Zhai, Xin Kong, Tingzhong Fu, Daniel F. N. Gordon, Xueli An, Benjamin Busam (Huawei Munich Research Center; Imperial College London; TUM)
flowchart LR
T[Manipulation task] --> S1[Stage 1: REACH]
S1 --> S2[Stage 2: GRASP]
S2 --> S3[Stage 3: TRANSPORT]
S3 --> S4[Stage 4: PLACE]
S1 -.dense reward.-> R[Stage-aligned reward signal]
S2 -.dense reward.-> R
S3 -.dense reward.-> R
S4 -.dense reward.-> R
R --> RL[STA-TPO offline preference<br/>+ STA-PPO online interaction]
Reach→Grasp→Transport→Place is the canonical pick-and-place decomposition; other tasks use stages such as Push/Pull, Lift, Upright, and Goal. The number of stages varies (typically 3–5) per task.
Existing RL fine-tuning treats long-horizon action trajectories as linguistic sequences and applies trajectory-level objectives — Trajectory-wise Preference Optimization (TPO) or PPO — over the whole trajectory. This gives coarse credit assignment and unstable training. The paper's key observation: unlike language, where meaning is preserved despite flexible word order, action trajectories progress through causally chained stages with different learning difficulties, motivating progressive stage-wise optimization.
STARE (Stage-Aware Reinforcement) is a module that decomposes a long-horizon trajectory into semantically meaningful stages and emits dense, interpretable, stage-aligned reward signals. Segmentation is rule-based / event-driven: stage boundaries are set by geometric and semantic thresholds (δ_k) on end-effector translation/orientation and object state (e.g., Reach→Grasp when the gripper contacts the object; Transport→Place when the object is within a distance margin of the goal) rather than arbitrary temporal cuts.
STARE is plugged into both optimizers:
- STA-TPO — offline, stage-wise preference learning.
- STA-PPO — online, intra-stage interaction.
These are composed in the IPI pipeline (Imitation → Preference → Interaction): SFT initialization, then STA-TPO offline preference, then STA-PPO online interaction. Base VLAs fine-tuned: OpenVLA-7B and π0.5_base.
Evaluated on SimplerEnv and ManiSkill3 (OpenVLA-7B backbone). The full IPI pipeline reaches state-of-the-art 98.0% on SimplerEnv and 96.4% on ManiSkill3.
| Method | SimplerEnv | ManiSkill3 |
|---|---|---|
| SFT | 41.7% | 15.0% |
| GRAPE | 43.9% | 18.0% |
| STA-TPO | 51.6% | 20.8% |
| RL4VLA | 92.5% | 70.5% |
| STA-PPO | 94.6% | 93.4% |
| IPI (full) | 98.0% | 96.4% |
Ablation: disabling STARE at the final high-precision stages (Place, Upright) causes >20% success drops, showing stage-aligned signal matters most where precision is critical.
Reframes VLA RL fine-tuning from trajectory-level to stage-level credit assignment, addressing the instability of treating actions as flat token sequences. The event-driven stage segmentation needs no hand-tuned dense reward per task, and the IPI ordering (preference before interaction) yields large gains over both pure offline (STA-TPO) and pure online (STA-PPO) RL alone.
- arXiv 2512.05107
- ICLR 2026 listing