ICLR 2026 Sparse Imagination - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 ยท Authors: Junha Chun, Youngjoon Jeong, Taesup Kim ยท arXiv: 2506.01392 ยท Category: World model for actions ยท Trend tag: efficient world-model planning.
flowchart LR
Img[224x224 RGB obs] --> Enc[DINO-ViT-S/16<br/>14x14 = 196 tokens]
Enc --> Drop[Token dropout<br/>keep 1-p fraction]
Drop --> WM[Transformer world model<br/>randomized grouped attention]
WM --> Roll[Sparse latent rollout<br/>imagination]
Roll --> Plan[Planning / trajectory<br/>optimization]
Plan --> Act[Action]
World-model-based planning lets agents simulate futures, but latent rollout over many visual tokens is expensive โ a severe constraint in resource-limited robotics where planning happens online.
A sparsely trained transformer visual world model that reduces the number of tokens processed during forward prediction:
-
Randomized grouped attention during training so the model can flexibly adjust the token count at inference based on available compute (process only
(1โp)NofNtokens). -
Sparse imagination: drop a fraction
pof patch tokens during latent rollout, accelerating planning while preserving control fidelity. - Encoder is DINO-ViT-S/16 on 224ร224 images, yielding 14ร14 = 196 spatial tokens. Applicable from test-time trajectory optimization to tasks driven by recent VLAs.
- Planning-time reductions at 50% dropout (per Table 2): PushT 173sโ82s (52.6%), Pointmaze 184sโ93s (49.5%), Wall 79sโ53s (32.9%), Block Pushing 297sโ208s (30.0%).
- Task success preserved or improved at moderate dropout (10โ50%): Wall 85.0%โ95.0%, Rope 63.3%โ73.3%, Pointmaze 98.3%โ100.0%, PushT 75.0%โ70.0%.
- Evaluated across six environments: Pointmaze, Wall, PushT, Granular, Rope, Block Pushing.
Shows that dropping visual tokens during imagination can nearly halve planning latency with little or no loss in task performance, making world-model planning more practical for compute-constrained robots. A targeted efficiency contribution within the "world model as actionable substrate" agenda.
- arXiv: https://arxiv.org/abs/2506.01392
- OpenReview: https://openreview.net/forum?id=faxcxKINBC
- ICLR 2026 Survey
- VLA Architectures review (world-model category)
โ Back to ICLR-2026