ICRA 2026 VLA Reasoner - Heungwoo/research GitHub Wiki

VLA-Reasoner — plug-and-play online MCTS that lets any VLA foresee futures before acting

Venue: ICRA 2026 · Authors: Wenkai Guo, Guanxing Lu, Haoyuan Deng, Zhenyu Wu, Yansong Tang, Ziwei Wang (NTU · Tsinghua SIGS · BUPT; Ziwei Wang corresponding) · arXiv: 2509.22643 Category: Reasoning-augmented VLA Trend tag: MCTS / test-time reasoning for action

Approach diagram

flowchart LR
  OBS["obs oₜ + instruction"] --> VLA["off-the-shelf VLA πθ"]
  VLA --> SEED["seed root: sample N actions"]
  SEED --> KDE["KDE → top-k candidates"]
  KDE --> WM["world model 𝒲<br/>(iVideoGPT 600M)<br/>oᵢ₊₁ = 𝒲(aᵢ, oᵢ)"]
  WM --> VAL["offline value/reward<br/>ResNet-34 + MLP"]
  VAL --> UCB["UCB backup over Q(oᵢ)"]
  UCB --> SEL["selected aᵗᴿᵉᵃˢᵒⁿᵉʳ"]
  VLA --> INJ["inject: aₜ = α·aᴠᴸᴬ + (1−α)·aᴿᵉᵃˢᵒⁿᵉʳ"]
  SEL --> INJ
  INJ --> ACT["executed action aₜ"]
Loading

Problem

VLAs trained by scaling imitation learning predict a short-sighted next action. On long-horizon trajectories small per-step errors accumulate, and the policy has no mechanism to look ahead and detect that a locally-plausible action leads to a dead end. VLA-Reasoner asks: can we add test-time lookahead to an already-trained VLA, without retraining the policy, by reasoning over imagined future states?

Method

VLA-Reasoner wraps a frozen, off-the-shelf VLA with online Monte Carlo Tree Search at inference time. The per-step VLA prediction seeds the search root.

  • Candidate sampling. From N actions sampled from πθ, Kernel Density Estimation selects the top-k closest (in L2) to the VLA proposal, focusing the tree on plausible actions while keeping exploration efficient in a large continuous action space.
  • World model rollout. An action-aware world model (iVideoGPT architecture, 600M params, finetuned on robot datasets plus failure demonstrations) generates future visual states oᵢ₊₁ = 𝒲(aᵢ, oᵢ), so each branch is rolled out into an imagined outcome — the actions are the "rationales."
  • Value / reward. An offline value estimator (ResNet-34 + 2-layer MLP) scores predicted future frames; it is trained with MSE against linearly-interpolated ground-truth rewards (0→1) on downsampled offline trajectories. Node values aggregate immediate reward and children via visit-count-weighted backup, where counts come from KDE density rather than explicit counting.
  • Selection. A UCB criterion balances exploitation of Q(oᵢ₊₁) against exploration.
  • Action injection. The executed action blends VLA and search outputs, aₜ = α·aᵗᴠᴸᴬ + (1−α)·aᵗᴿᵉᵃˢᵒⁿᵉʳ (best α = 0.6).

The search runs online (per step at inference); the value/reward models are trained offline. The wrapper is plug-and-play and attaches to any VLA policy.

Results

LIBERO simulation (success rate, baseline → +VLA-Reasoner):

  • OpenVLA-SFT: 76.0% → 81.0% (+5.0 pp)
  • Octo-Small: 26.5% → 37.3% (+10.8 pp)
  • SpatialVLA: 34.0% → 41.8% (+7.8 pp)

Real-world (5 tasks, 20 trials each):

  • OpenVLA: 22% → 41% (+19 pp)
  • π₀-FAST: 64% → 74% (+10 pp)

Gains are consistent across tasks, environments, and robot embodiments.

Significance

VLA-Reasoner is a clean test-time-scaling recipe for manipulation: keep the VLA frozen, add a world model to imagine futures and an offline value head to score them, and search with MCTS. Because it is policy-agnostic it complements both small (Octo) and large (OpenVLA, π₀-FAST) backbones, making lookahead reasoning a drop-in upgrade rather than a new architecture.

Links

Related pages

← Back to ICRA-2026

⚠️ **GitHub.com Fallback** ⚠️