NeurIPS 2025 VLA OS - Heungwoo/research GitHub Wiki

VLA-OS — Structuring and Dissecting Planning Representations

Venue: NeurIPS 2025 (poster) · Authors: Chongkai Gao, … Lin Shao (NUS Lins Lab; collaborators USTC, Tsinghua, NTU) · arXiv: 2506.17561 Category: VLA Architecture — controlled study

Approach diagram

flowchart LR
  subgraph Variants[VLA-OS variants]
    OSA[OS-A: Action-Only<br/>no planner]
    OSI[OS-I: Integrated<br/>planner + action one stream]
    OSH[OS-H: Hierarchical<br/>planner → action]
  end
  Variants --> Bench[Controlled benchmark:<br/>2D + 3D · rigid + deformable ·<br/>gripper + dex hand]
  Bench --> F[Finding 1: visual-grounded planning > language planning]
  Bench --> F2[Finding 2: Hierarchical &gt; Integrated on generalization/planning<br/>both &gt; Action-Only]
Loading

Problem

The VLA field has been exploring planning representations (language subtasks, subgoal images, action CoT, latent plans) piecemeal — each paper invents one and claims it helps, but nobody has controlled the architecture to isolate "which planning representation" from "which training data / backbone / task."

Method

A unified VLA architecture suite (OS-A, OS-I, OS-H) that differs only in how planning is exposed:

  • OS-A (Action-Only): no explicit planner; VLA directly emits actions.
  • OS-I (Integrated): planner + action share one decoding stream.
  • OS-H (Hierarchical): explicit planning stage feeds action generation.

All three variants share backbone, data mix, and training recipe. Ablations run across 2D vs. 3D vision, rigid vs. deformable objects, and gripper vs. dexterous hand.

Results

  • Visual-grounded planning beats language-grounded planning (subgoal images > subtask text) — and is faster/cheaper: image-foresight heads need ~7 forward passes vs. hundreds for autoregressive language plans (~100× efficiency), since language plans cost ~2000 planning tokens vs. 8 action tokens (accumulating errors).
  • Hierarchical-VLA shows superior generalization and task-planning over Integrated-VLA (overall task performance is comparable), and both add planning over Action-Only — but Hierarchical pays a cost in slower training/inference.
  • For ~5,000-demo task scale, the optimal LLM backbone is ~0.5B params (total model <1B).

Significance

The 2025–2026 hierarchical-VLA design (π0.5 / π0.7 / GR00T N1.6) is vindicated by VLA-OS's controlled study. The finding that visual-grounded > language-grounded also directly motivates π0.7's BAGEL-generated subgoal images and Unified Diffusion VLA's joint image+action denoising.

Links

Related pages

← Back to NeurIPS-2025

⚠️ **GitHub.com Fallback** ⚠️