CVPR 2026 CoWVLA - Heungwoo/research GitHub Wiki
Venue: CVPR 2026 Category: Goal-Image / world-model-as-prompt VLA Trend tag: Trend 3 Affiliations: HIT + Li Auto + BAAI + UNSW + PKU
flowchart LR
OBS["current obs"] --> VAE["pretrained video VAE"]
VAE --> S["structure latent"]
VAE --> M["motion latent"]
S --> COW["Chain of World<br/>predict motion chain"]
M --> COW
COW --> TERMS["terminal latent frame"]
TERMS --> POL["VLA policy"]
OBS --> POL
POL --> ACT["action"]
π0.7 introduces subgoal-image-as-prompt with BAGEL-14B but operates in pixel space (a 14 B generator). Pixel-space is expensive and lossy — most of the bits encode appearance, not motion. Can the same benefit be had in a latent motion space?
Use a pretrained video-VAE (VidTwin, from Microsoft, fine-tuned on ~237k robot-manipulation videos) to factorize clips into structure latents (appearance) and motion latents (dynamics). The VLA "thinks" by predicting a chain of motion latents culminating in a terminal latent frame — i.e., a sequence of latent subgoals rather than pixel-space subgoals. The backbone is an 8.5 B-parameter Emu3 autoregressive VLM: during pre-training it infers a continuous latent motion chain + terminal frame from an instruction and initial frame; during co-fine-tuning the latent dynamics are aligned with discrete action prediction by jointly modeling sparse keyframes and action sequences in a unified autoregressive decoder.
On robotic simulation benchmarks CoWVLA outperforms world-model (UniVLA, FlowVLA, CoT-VLA, WorldVLA), latent-action (LAPA, villa-X, TLA) and standard VLA (OpenVLA, π₀, GR00T-N1) baselines:
| Benchmark | Result |
|---|---|
| LIBERO (avg of 4 suites) | 95.6% (Spatial 97.2 / Object 97.8 / Goal 94.6 / Long 92.8) |
| SimplerEnv-WidowX (avg of 4) | 76.0% |
| SimplerEnv-Google Robot (avg) | 60.9% |
| CALVIN ABCD→D / ABC→D (avg success length) | 4.473 / 4.211 |
The paper claims "moderate computational efficiency" — slower/heavier than latent-action methods like LAPA but cheaper than world-model methods like UniVLA. Real-robot validation is limited: cup-grasping on a Realman RM75B (127 episodes) across lighting conditions.
CoWVLA is the latent-space dual of π0.7's pixel-space subgoals. Both bet that future-state prediction helps action prediction; CoWVLA bets the "thinking" can be done in a compact disentangled video-VAE latent (structure vs. motion) rather than full pixel space, avoiding capacity wasted on background reconstruction. Sits cleanly in cluster E of the Goal-Image Conditioning review (foundation-model-as-prompt) but routes the "prompt" through a latent space.
- arXiv: 2603.03195
- Project:
fx-hit.github.io/cowvla-io
← Back to CVPR-2026