ICLR 2026 Sparse Imagination - Heungwoo/research GitHub Wiki

Sparse Imagination โ€” token-sparse rollouts for efficient visual world-model planning

Venue: ICLR 2026 ยท Authors: Junha Chun, Youngjoon Jeong, Taesup Kim ยท arXiv: 2506.01392 ยท Category: World model for actions ยท Trend tag: efficient world-model planning.

Approach diagram

flowchart LR
  Img[224x224 RGB obs] --> Enc[DINO-ViT-S/16<br/>14x14 = 196 tokens]
  Enc --> Drop[Token dropout<br/>keep 1-p fraction]
  Drop --> WM[Transformer world model<br/>randomized grouped attention]
  WM --> Roll[Sparse latent rollout<br/>imagination]
  Roll --> Plan[Planning / trajectory<br/>optimization]
  Plan --> Act[Action]
Loading

Problem

World-model-based planning lets agents simulate futures, but latent rollout over many visual tokens is expensive โ€” a severe constraint in resource-limited robotics where planning happens online.

Method

A sparsely trained transformer visual world model that reduces the number of tokens processed during forward prediction:

  • Randomized grouped attention during training so the model can flexibly adjust the token count at inference based on available compute (process only (1โˆ’p)N of N tokens).
  • Sparse imagination: drop a fraction p of patch tokens during latent rollout, accelerating planning while preserving control fidelity.
  • Encoder is DINO-ViT-S/16 on 224ร—224 images, yielding 14ร—14 = 196 spatial tokens. Applicable from test-time trajectory optimization to tasks driven by recent VLAs.

Results

  • Planning-time reductions at 50% dropout (per Table 2): PushT 173sโ†’82s (52.6%), Pointmaze 184sโ†’93s (49.5%), Wall 79sโ†’53s (32.9%), Block Pushing 297sโ†’208s (30.0%).
  • Task success preserved or improved at moderate dropout (10โ€“50%): Wall 85.0%โ†’95.0%, Rope 63.3%โ†’73.3%, Pointmaze 98.3%โ†’100.0%, PushT 75.0%โ†’70.0%.
  • Evaluated across six environments: Pointmaze, Wall, PushT, Granular, Rope, Block Pushing.

Significance

Shows that dropping visual tokens during imagination can nearly halve planning latency with little or no loss in task performance, making world-model planning more practical for compute-constrained robots. A targeted efficiency contribution within the "world model as actionable substrate" agenda.

Links

Related pages

โ† Back to ICLR-2026

โš ๏ธ **GitHub.com Fallback** โš ๏ธ