CoRL 2026 SG WAM - Heungwoo/research GitHub Wiki

CoRL 2026 โ€” SG-WAM: Self-Guided World Modeling in Geometry-Aware Policy Space

Venue: CoRL 2026 (Austin, TX, Nov 9โ€“12). Paper: arXiv 2608.01397. Representative of: latent/representation-space World-Action Models โ€” predict future dynamics in a geometry-aware policy space instead of pixels. Companions: World Models ยท VLA Hybrid Architectures ยท CoRL 2026 survey.

SG-WAM overview: world modeling in a geometry-aware, policy-derived latent space (figure from the authors, arXiv 2608.01397, ยฉ the authors)

1. Problem

World-Action Models (WAMs) want a policy that also predicts how actions change the scene. Two failure modes recur: pixel/observation-space prediction spends capacity on perceptual detail irrelevant to control, while latent WAMs often regress toward auxiliary targets that are mismatched with the action-generation representation and carry no geometric grounding. Manipulation, however, needs to know where and how an action reshapes the scene. SG-WAM asks for a prediction space that is at once aligned with action generation and sufficiently geometry-aware.

2. Method

Three jointly-trained components:

  • Geometry-aware policy states. A frozen VGGT foundation model shapes the VLM's visual tokens via cosine-similarity alignment, injecting spatial structure into the policy representation (a teacher for shape, not for future targets).
  • Self-Guided World Predictor (SGWP). A small set of learnable dynamics tokens forecasts future latent states conditioned on the intervening actions. Both current and future representations come from the same policy backbone โ€” current states from the online VLM, future targets from an EMA copy โ€” which removes the target/policy mismatch that plagues auxiliary-target latent WAMs.
  • Conditional flow-matching action expert. Generates 8-step action chunks from the unified policy context, keeping the dynamics tokens on the action-generation pathway.

Backbone is a Qwen3.5-0.8B VLM, ~0.9B parameters total, with no large-scale embodied pretraining; 8 dynamics tokens; trained ~40k steps on LIBERO. The geometry teacher, SGWP and EMA pathway are auxiliary and removed at inference.

3. Results

  • LIBERO (4 suites): 98.5% average success.
  • LIBERO-Plus (zero-shot OOD): 73%.
  • Real-world (in-distribution): Pick-and-Place 75%, Towel-Folding 45%, Toolbox-Organization 50%; SG-WAM leads baselines under both ID and OOD (background shift, lighting change, novel object) conditions.
  • Ablations: removing world modeling โˆ’1.9 pp, removing geometric supervision โˆ’0.9 pp average; 8 dynamics tokens optimal (16 tokens dropped to 97.2%).

4. Why it matters

It reframes the world model not as a pixel predictor but as a self-supervised objective inside the policy's own geometry-aware latent space, so the same tokens that predict dynamics also drive action generation. That yields strong LIBERO and OOD numbers at 0.9B without embodied pretraining โ€” evidence that representation choice, not scale, can carry WAM performance.

Limitations (reviewer): no explicit limitations section; empirically, novel-object real-world grasping drops sharply (40% vs 75% ID), and long-horizon LIBERO-Long (96.2%) trails the Object suite (99.8%) โ€” the geometry teacher (VGGT) is also an external dependency.

5. Links

โ† Back to CoRL 2026 survey ยท Home