RSS 2026 Act2Goal - Heungwoo/research GitHub Wiki

Act2Goal: From World Model To General Goal-conditioned Policy

Venue: RSS 2026 (Sydney, Jul 13–17) · Session: World Models & Memory · paper #15 Authors: Pengfei Zhou, Liliang Chen, Shengcong Chen, Di Chen, Wenzhi Zhao, Rongjun Jin, Guanghui Ren, Jianlan Luo arXiv: 2512.23541 · program page

Summary compiled from the arXiv paper (v1); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Act2Goal method overview (Figure 1 of arXiv 2512.23541, © the authors)

The robot receives a visual goal (left, table with flowers in a vase), imagines a sequence of intermediate visual states toward it via a goal-conditioned world model (thought cloud, "Plan"), and executes the planned actions in the real world ("Action") — the core imagine-then-act loop of Act2Goal.

Problem

Visual goals are a compact, unambiguous alternative to language for specifying manipulation tasks, but existing goal-conditioned policies predict actions in a single step without modeling task progress, so they degrade on long-horizon tasks and overfit demonstrated state–action mappings. The authors (Agibot Research) ask how a policy can explicitly reason about the visual dynamics needed to reach a distant goal.

Method

Act2Goal couples a Goal-Conditioned World Model (GCWM) — built on the Genie Envisioner architecture with language conditioning removed and a goal image concatenated to the observation — with an isomorphic flow-matching action expert (1.6B-parameter Video DiT + 160M action DiT, 28 blocks each, joined by cross-attention). Multi-Scale Temporal Hashing (MSTH) splits the imagined trajectory into dense proximal frames (stride-r sampled up to horizon P, with actions at every timestep) for closed-loop control and logarithmically spaced distal frames that anchor global consistency; only proximal actions are executed. Training is two-stage offline imitation (world-model fine-tuning, then joint flow-matching of video and actions), plus optional reward-free online improvement: HER-style hindsight goal relabeling of self-collected rollouts with LoRA-only updates on the edge device.

Results

On four Robotwin 2.0 simulation tasks, Act2Goal beats DP-GC, π0.5-GC, and HyperGoalNet on all Easy-mode tasks (e.g., Pick Bottles 0.80 vs. 0.13 for π0.5-GC) and 3 of 4 Hard-mode tasks. On an AgiBot Genie-01 robot across three real tasks, it scores ID/OOD 0.93/0.90 (whiteboard word writing), 0.75/0.48 (dessert plating), 0.45/0.30 (plug-in), while DP-GC and HyperGoalNet are near 0. Online autonomous improvement converges in ~3 rounds with up to 8× success-rate gains in simulation; on the real OOD plug-in task success climbs from 0.30 to 0.90, and even failed-only rollouts yield improvement.

Significance

First integration of a world model into goal-conditioned policy learning per the authors, and a practical recipe for reward-free on-robot self-improvement (HER + LoRA) that needs no human labels. Fits the wiki's Review-World-Models thread and the deployment-time-adaptation discussion in RL.

← Back to RSS 2026 survey · RSS-2026-Papers · Home