IROS 2026 AtomVLA - Heungwoo/research GitHub Wiki

IROS 2026 — AtomVLA: Scalable Post-Training for Manipulation via Predictive Latent World Models

Venue: IROS 2026 (Pittsburgh) · paper #33 · Huazhong Univ. of S&T · HKU · Tsinghua · Beihang (Sun, Xu, Cao, … Chen). Paper: arXiv 2603.08519 (Mar 2026) · datasets/checkpoints/code to be released. The WAM×VLA datapoint of IROS 2026 — a latent world model used to score a VLA's action chunks during offline post-training, closing the instruction-grounding gap on long-horizon tasks. Companions: World Models · VLA Hybrid Architectures · Multi-Task VLA · DYNA-2 · IROS 2026 survey.

AtomVLA's two-stage pipeline — Stage I SFT (Qwen3-VL backbone + flow-matching action head); Stage II post-training: the frozen VLA rolls out k candidate action chunks in the environment, a World-Model encoder scores them against subtask/goal frames, and the resulting reward drives offline GRPO on the action head (architecture figure from Sun et al., arXiv 2603.08519, © the authors)

1. Problem

VLAs execute complex multi-step behaviors better with robust instruction grounding, but current SFT relies on coarse high-level instructions — no explicit intermediate guidance — so long-horizon tasks suffer compounding errors. The paper frames this as an instruction-grounding gap and seeks a scalable post-training recipe to close it.

2. Method

AtomVLA is "the first subtask-aware VLA framework integrated with a scalable offline post-training pipeline."

  • Atomic-subtask decomposition — an LLM decomposes high-level demonstrations into fine-grained atomic subtasks (the intermediate guidance the SFT signal lacks).
  • Latent world-model scoring — a pretrained predictive world model scores candidate action chunks against subtask goals in latent space, mitigating error accumulation and improving long-horizon robustness. (This is the reconstruction-free, deployable latent-WAM flavor — cf. ω-0/DYNA-2 — used as an evaluator/critic, not an in-path generator.)
  • Offline GRPO — the latent scoring enables Group Relative Policy Optimization without online physical-robot rollouts (the expensive part of RL-for-VLA — cf. DyGRO-VLA, RL for VLA).

3. Results

  • LIBERO 97.0%, LIBERO-PRO 48.0% avg success vs baselines; strong robustness under perturbations.
  • Real-world on the Galaxea R1 Lite — broad applicability, especially long-horizon tasks.
  • Datasets / checkpoints / code to be publicly released.

4. Why it matters (WAM lens)

AtomVLA is a clean instance of IROS 2026's WAM trend (survey §5.1): the world model has moved from pixel generator to latent evaluator inside a VLA post-training loop. It combines three threads the wiki tracks — latent WAM (World Models), subtask/instruction grounding (Multi-Task VLA, AtomVLA's "instruction gap" ≈ LangForce/DISC's language-shortcut problem), and offline RL for VLA (DyGRO-VLA). The offline-GRPO-via-latent-scoring move is the notable systems contribution: RL-quality post-training without robot rollouts.

Limitations (reviewer): results are LIBERO/LIBERO-PRO sim + one real platform; the world model's scoring fidelity is the ceiling (a bad latent critic misranks chunks); LLM subtask decomposition quality is an unmeasured dependency.

5. Links

← Back to IROS 2026 survey · Home