RSS 2026 Self Improving Robot Policy with Compositional - Heungwoo/research GitHub Wiki

Self-Improving Robot Policy with Compositional World Model

Venue: RSS 2026 (Sydney, Jul 13–17) · Session: World Models & Memory · paper #12 Authors: Jiazhi Yang, Kunyang Lin, Wencong Zhang, Jinwei Li, Tianwei Lin, Longyan Wu, Ya-Qin Zhang, Hao Zhao, Ping Luo, Zhizhong Su, Hongyang Li, Xiangyu Yue, Li Chen arXiv: 2602.11075 · program page

Summary compiled from the arXiv paper (v2, "RISE: Self-Improving Robot Policy with Compositional World Model"); all numbers quoted from the paper. Trend context: RSS 2026 survey.

RISE framework overview (Figure 1 of arXiv 2602.11075, © the authors)

Panel (a) shows why physical-world RL is costly: laborious resets, monitoring, and slow serial execution. Panel (b) is the core idea — a Compositional World Model whose dynamics model imagines multi-view futures for candidate actions and whose value model scores them (low vs high advantage), driving online RL entirely in imaginary space. Panel (c) shows the three real-world evaluation tasks with success-rate bars versus RECAP (+35% brick sorting, +45% backpack packing, +35% box closing).

Problem

VLA models trained by imitation remain brittle in contact-rich and dynamic manipulation because small execution deviations compound (exposure bias), and on-policy RL in the physical world is blocked by safety risk, hardware cost, and manual environment resets. RISE asks whether the RL loop can be moved entirely into a learned "imagination" environment.

Method

RISE pairs a controllable dynamics model — initialized from the GE-base variant of Genie Envisioner and trained with a Task-centric Batching strategy for action controllability — with a progress value model initialized from the pre-trained π0.5 VLA and trained with progress-estimate plus Temporal-Difference objectives on both success and failure data. The dynamics model synthesizes multi-view futures roughly 300x faster than the cited alternative (which needs ~minutes for 25 multi-view observations); the value model converts imagined outcomes into chunk-wise advantages. The policy (fine-tuned π0.5, following the RECAP recipe) is trained advantage-conditioned via flow matching: it rolls out in imagination, gets its actions scored into N discrete advantage bins, and is updated with EMA blending in a closed self-improving loop — no physical interaction during RL.

Results

On a dual 7-DoF AgileX bimanual robot across three long-horizon tasks, RISE reaches 85% success on Dynamic Brick Sorting, 85% on Backpack Packing, and 95% on Box Closing, versus RECAP's 50/40/60% and π0.5's 35/30/35%; π0.5+DAgger, π0.5+PPO, and π0.5+DSRL all trail further (e.g., 10-15% on brick sorting). That is the abstract's ">+35% / +45% / +35% absolute" improvement over prior art; scores (0-10 scale) rise correspondingly (9.78/9.50/9.88).

Significance

A concrete demonstration that world-model "imagination RL" can beat physical-data RL pipelines (RECAP, DSRL) on real robots — a strong entry for the Review-World-Models thread and for the broader question of scalable self-improvement beyond imitation (RL).

← Back to RSS 2026 survey · RSS-2026-Papers · Home