RSS 2026 Collaborating Visual and Parameter Spaces - Heungwoo/research GitHub Wiki

Collaborating Visual and Parameter Spaces for Consistent Long-Horizon Embodied World Model

Venue: RSS 2026 (Sydney, Jul 13–17) · Session: World Models & Memory · paper #14 Authors: Longyu Chen, Heng Li, Wei Yang, Manqi Zhao, Dongsheng Jiang arXiv: 2606.28804 · program page

Summary compiled from the arXiv paper (v1, titled "ViPSim: Collaborating Visual and Parameter Spaces for Consistent Long-Horizon Embodied World Models"); all numbers quoted from the paper. Trend context: RSS 2026 survey.

ViPSim qualitative rollouts (Figure 1 of arXiv 2606.28804, © the authors)

Two qualitative comparisons against the EnerVerse-AC baseline. Top: a long-horizon deformable-object task ("fold green shorts") where ViPSim keeps the garment's structure intact out to chunk 62 while the baseline degrades. Bottom: cross-embodiment generalization — trained solely on AgiBot data, the model emulates Droid-robot actions in an out-of-distribution scene ("grasp the sponge scrubber").

Problem

Generative Embodied World Models (EWMs) predict future visual states from observations and actions, offering a risk-free way to evaluate VLA policies. But a representation gap between low-dimensional actions and high-dimensional video synthesis breaks geometric correspondence, causing accumulated trajectory drift and inconsistent robot-object interactions over long-horizon rollouts.

Method

ViPSim couples two conditioning spaces around a chunk-based autoregressive video diffusion backbone (L = 16 frames per chunk at 320×512). The Visual Space supplies pixel-aligned structural priors: rendered end-effector action maps (50-pixel-radius circles, color-coded per arm), per-pixel Plücker camera embeddings, Video-Depth-Anything depth maps, and robot morphology masks from a fine-tuned detector plus SAM2. The Parameter Space injects precise numerical drivers — quaternion-reparameterized 8-D per-arm actions (16-D dual-arm) and flattened 12-D camera matrices — through a lightweight Physics Encoder into cross-attention. The design is backbone-agnostic, demonstrated on both a UNet backbone (DynamiCrafter, 40K iterations) and a DiT backbone (Wan2.2-TI2V-5B, 20K iterations, flow matching), trained on 8×80GB GPUs; end-to-end inference runs at 1.918 s per chunk.

Results

On 10 AgiBotWorld-Beta tasks (1,000 training trajectories, 94 unseen test clips), ViPSim(DiT) reaches PSNR 20.35 / SSIM 0.80 / LPIPS 0.19 versus 17.93 / 0.73 / 0.25 for EnerVerse-AC. On the EWMBench protocol its overall score is 5.5697 versus 5.0386 for EnerVerse-AC, with motion-correctness DYN improving from 0.5723 to 0.7534. Action-swap tests show generated motion follows transplanted numerical actions regardless of scene, and ablations confirm each visual/parameter component contributes (full model LPIPS 0.2356 vs 0.2571 for action-map only).

Significance

A precision-centric recipe for making world models trustworthy as VLA evaluators: dense pixel-aligned grounding plus raw numerical conditioning suppresses the drift that plagues action-conditioned video generation, including on deformable objects. Related wiki threads: Review-World-Models · Review-VLA-Memory.

← Back to RSS 2026 survey · RSS-2026-Papers · Home