RSS 2026 Causal World Modeling for Robot - Heungwoo/research GitHub Wiki

Causal World Modeling for Robot Control

Venue: RSS 2026 (Sydney, Jul 13–17) · Session: World Models & Memory · paper #16 Authors: Lin Li, Qihang Zhang, Yiming Luo, Shuai Yang, Ruilin Wang, Zhangluyao, Mingrui Yu, Zelin Gao, Nan Xue, Boyu Zhou, Xing Zhu, Mingyu Ding, Yujun Shen, Yinghao Xu arXiv: 2601.21998 · program page

Summary compiled from the arXiv paper (v2, Ant Group / Robbyant); all numbers quoted from the paper. Note: the RSS program abstract names the model CauVA, while the arXiv version calls it LingBot-VA — same framework. Trend context: RSS 2026 survey.

LingBot-VA causal video-action policy overview (Figure 1 of arXiv 2601.21998, © the authors)

Figure 1 shows the full system: internet and robot videos pretrain a causal video-action policy in which a language model, an autoregressive video model, and an action model share one sequence ("future imagination" + robot action outputs); the bottom-left panel illustrates the world-model / inverse-dynamics / execution loop over consecutive observations, and the bottom-right bar charts summarize gains over π0.5 in real-world success/progress, simulation (RoboTwin, LIBERO), long temporal memory, and data efficiency.

Problem

Feedforward VLAs entangle scene understanding, physical dynamics, and motor control in one network trained from a single supervision signal, hurting sample efficiency and generalization. Existing world-model attempts (interactive simulators, chunk-based video-action diffusion, offline subgoal video generators) suffer a reactivity gap (open-loop rollouts ignore real-time feedback), limited long-term memory across chunks, and non-causal bidirectional attention within segments.

Method

LingBot-VA is an autoregressive diffusion world model that interleaves video latents and action tokens in a single causal sequence and generates chunks of K frames (with τ = 4 actions per frame) via conditional flow matching. Components: (1) a Mixture-of-Transformers dual-stream architecture — a video stream initialized from Wan2.2-5B (d_v = 3072, 30 layers) plus a narrower action stream (d_a = 768, ~350M extra params; 5.3B total) fused by cross-modal attention; (2) a closed-loop rollout with persistent KV-cache memory over the whole interleaved trajectory, re-grounded each step by real observations; (3) an asynchronous inference pipeline that overlaps prediction with execution, using a Forward-Dynamics-Model (FDM) grounding step to replace stale visual forecasts with feedback-conditioned predictions, plus noisy-history augmentation enabling partial denoising for fast action decoding. Actions decode via an inverse-dynamics model conditioned on predicted future latents and full observation/action history; pretraining uses ~16K hours of data (Agibot, RoboMind, InternData-A1, OXE, UMI data, RoboCOIN).

Results

On RoboTwin 2.0 (50 bimanual tasks), LingBot-VA averages 92.93% (Easy) / 91.55% (Hard), +4.2/+4.6 points over the second-best method (Motus) and ahead of π0.5 (82.7/76.8) and π0 — with the largest margins on horizon-3 tasks (+8.2 Easy / +9.1 Hard). On LIBERO it reaches a 98.5% average (LIBERO-Object 99.6%, LIBERO-Long 98.5%, LIBERO-Spatial 98.5%), a new state of the art among the compared foundation VLAs. On six real-world tasks (long-horizon, precision, deformable; 50 demos each) it beats π0.5 on every task — Figure 1 reports aggregate 59.2 vs 39.2 success and 79.2 vs 65.4 progress score — and memory probes (Wipe Plate, Search Box) hit 100% vs π0.5's ~47–50%. With only 10 demos it shows +15.6% progress on "Make Breakfast" and +10.3% on RoboTwin 2.0 Easy over π0.5.

Significance

A strong data point for the claim that causal video world modeling is a second pretraining foundation alongside vision-language pretraining, and its KV-cache persistent context directly addresses the memory theme of this session — see Review-World-Models and Review-VLA-Memory. Code, checkpoints, and models are public.

← Back to RSS 2026 survey · RSS-2026-Papers · Home