IROS 2026 AtomVLA - Heungwoo/research GitHub Wiki
IROS 2026 — AtomVLA: Scalable Post-Training for Manipulation via Predictive Latent World Models
Venue: IROS 2026 (Pittsburgh) · paper #33 · Huazhong Univ. of S&T · HKU · Tsinghua · Beihang (Sun, Xu, Cao, … Chen). Paper: arXiv 2603.08519 (Mar 2026) · datasets/checkpoints/code to be released. The WAM×VLA datapoint of IROS 2026 — a latent world model used to score a VLA's action chunks during offline post-training, closing the instruction-grounding gap on long-horizon tasks. Companions: World Models · VLA Hybrid Architectures · Multi-Task VLA · DYNA-2 · IROS 2026 survey.

1. Problem
VLAs execute complex multi-step behaviors better with robust instruction grounding, but current SFT relies on coarse high-level instructions — no explicit intermediate guidance — so long-horizon tasks suffer compounding errors. The paper frames this as an instruction-grounding gap and seeks a scalable post-training recipe to close it.
2. Method
AtomVLA is "the first subtask-aware VLA framework integrated with a scalable offline post-training pipeline."
- Atomic-subtask decomposition — an LLM decomposes high-level demonstrations into fine-grained atomic subtasks (the intermediate guidance the SFT signal lacks).
- Latent world-model scoring — a pretrained predictive world model scores candidate action chunks against subtask goals in latent space, mitigating error accumulation and improving long-horizon robustness. (This is the reconstruction-free, deployable latent-WAM flavor — cf. ω-0/DYNA-2 — used as an evaluator/critic, not an in-path generator.)
- Offline GRPO — the latent scoring enables Group Relative Policy Optimization without online physical-robot rollouts (the expensive part of RL-for-VLA — cf. DyGRO-VLA, RL for VLA).
3. Results
- LIBERO 97.0%, LIBERO-PRO 48.0% avg success vs baselines; strong robustness under perturbations.
- Real-world on the Galaxea R1 Lite — broad applicability, especially long-horizon tasks.
- Datasets / checkpoints / code to be publicly released.
4. Why it matters (WAM lens)
AtomVLA is a clean instance of IROS 2026's WAM trend (survey §5.1): the world model has moved from pixel generator to latent evaluator inside a VLA post-training loop. It combines three threads the wiki tracks — latent WAM (World Models), subtask/instruction grounding (Multi-Task VLA, AtomVLA's "instruction gap" ≈ LangForce/DISC's language-shortcut problem), and offline RL for VLA (DyGRO-VLA). The offline-GRPO-via-latent-scoring move is the notable systems contribution: RL-quality post-training without robot rollouts.
Limitations (reviewer): results are LIBERO/LIBERO-PRO sim + one real platform; the world model's scoring fidelity is the ceiling (a bad latent critic misranks chunks); LLM subtask decomposition quality is an unmeasured dependency.
5. Links
- Official program: IROS 2026 (paper #33) · survey: IROS 2026
- Related: World Models · VLA Hybrid Architectures · DyGRO-VLA · Multi-Task VLA · ω-0
← Back to IROS 2026 survey · Home