ICLR 2026 Vlaser - Heungwoo/research GitHub Wiki

Vlaser — VLA with Synergistic Embodied Reasoning

Venue: ICLR 2026 Authors: Shanghai AI Lab + USTC + SJTU + Zhejiang + Nanjing + Fudan + Tsinghua + NUS + Northeastern + Shenzhen U. (Ganlin Yang, Tianyi Zhang, Haoran Hao, Weiyun Wang, Yibin Liu, … Yao Mu, Zhi Hou) Category: VLA Architecture — Reasoning-augmented · VLM-init study Trend tag: Embodied reasoning · VLM→VLA initialization · domain-shift mitigation

Approach diagram

flowchart LR
  Net[InternVL3 backbone<br/>2B / 8B] --> SFT[Stage 1: VLM SFT<br/>Vlaser-6M dataset]
  D6M[Vlaser-6M:<br/>1.5M grounding · 1.2M RoboVQA<br/>0.5M spatial · 0.4M planning<br/>2M in-domain SimplerEnv + RoboTwin2.0] --> SFT
  SFT --> VLM[Vlaser VLM<br/>SOTA on 12 embodied benchmarks]
  VLM --> AE[Action Expert<br/>flow-matching, π0-style<br/>shared self-attention<br/>non-causal VLA stream]
  Obs[image + lang + robot state q_t] --> AE
  AE --> A[Action chunk A_t<br/>H=4, 10 inference steps Euler]
Loading

Problem

VLAs are typically built by fine-tuning generic web-pretrained VLMs (Kim '24 OpenVLA, Black '24 π₀, Qu '25 SpatialVLA), but few studies examine how the VLM initialization choice affects downstream VLA. The hidden assumption is that reasoning-stronger VLMs make stronger VLAs — Vlaser tests it. They find a counter-intuitive result: out-of-domain embodied reasoning data does NOT translate cleanly to closed-loop manipulation; what helps is in-domain data from the same simulator/embodiment.

Detailed Method

Architecture

  • VLM backbone: InternVL3-2B (InternViT + Qwen2.5-1.5B) and InternVL3-8B (InternViT + Qwen2.5-7B). Two model sizes for compute-vs-quality trade-offs.
  • Action expert: π₀-style mixture-of-experts head — original VLM weights handle image/text streams; a separate weight set handles robot state + noised actions. Self-attention is shared between language and action experts; non-causal attention for the VLA stream.
  • Robot state q_t encoded as a state token; noised actions A_t^τ as action tokens; both projected via linear layer into the same embedding space as image/text features.

Vlaser-6M Dataset (the cornerstone)

Six categories, 6M total samples:

Category Size Sources
Embodied grounding (bbox + center points, normalized [0,1000]) 1.5M (+0.3M from SA-1B masks) RoboPoint, ShareRobot, Pixmo-Points, Paco-LaVIS, RefSpatial
RoboVQA (general state perception) 1.2M RoboVQA, Robo2VLM, RoboPoint, RefSpatial, OWMM-Agent
Spatial intelligence 0.5M (+0.1M manually annotated) SPAR, SpaceR-151k, VILASR; ScanNet, ScanNet++, CA-1M, ARKitScenes
Planning 0.4M Alpaca-15k, MuEP, WAP, LLaRP/Habitat, EgoPlan-IT, EgoCOT
In-domain VLA data (SimplerEnv Google Robot + WidowX, and RoboTwin2.0 Aloha-AgileX) 2M Synthesized QA on same observation space as eval; LLM-as-judge (Qwen2.5VL-32B) filtering drops ~10% lowest-scored
Total ~6M

The "in-domain 2M" subset is the key empirical lever — synthesized QA annotations directly on robot observation footage from the SimplerEnv simulator (WidowX + Google Robot) and RoboTwin2.0 (Aloha-AgileX), in the reasoning categories above.

Training Recipe (two-stage)

Stage 1 — Vision-Language Pretraining (SFT on Vlaser-6M):

  • Auto-regressive language-modeling loss L_lm = -log p(t_N | F_v(x; θ_v), F_t(y), t_{0:N-1}; Θ).
  • Initialized from InternVL3-2B / InternVL3-8B; full-parameter SFT.

Stage 2 — VLA fine-tuning (flow matching):

  • Action chunk A_t = [a_t, …, a_{t+H-1}] with horizon H = 4 (Action Chunk length = 4, Table 7).
  • Noisy action chunk A_t^τ = τA_t + (1-τ)ε, ε ~ N(0, I).
  • Loss L_vla = E_{p(A_t|o_t)} ‖v_θ(A_t^τ, o_t) - u(A_t^τ | A_t)‖² where u = ε - A_t.
  • Inference: forward Euler numerical integration, 10 inference steps; observation history length = 1 (single-frame). At eval only 2 of the 4 predicted actions are executed open-loop before re-planning (Execute Action length = 2, Table 7 / eval appendix) — note the Table 5 ablation separately reports a default "execute length H=4", an internal terminology clash in the paper.
  • Eval on SimplerEnv (WidowX Bridge + Google Robot) and RoboTwin 2.0 (Aloha-AgileX bimanual, Table 4). No real-robot experiments.

Training hyperparameters (Appendix Tables 6–7): VLM SFT — AdamW, peak LR 2e-5, cosine decay, 1 epoch / 5,000 steps, global batch 128, 150 warm-up steps, bf16, LLM seq len 16,384, dynamic resolution (patch 448, max 12 patches), all params trainable. VLA fine-tuning — AdamW, VLM & action-expert peak LR 5e-5, global batch 1,024, 10 epochs, seq len 384, all params trainable; WidowX checkpoint at 45,390 iters, Google Robot at 36,970 iters. (GPU hours not reported.)

Comprehensive Results

Embodied reasoning across 12 benchmarks (Table 1, normalized avg)

ERQA EgoPlan2 Where2place Pointarena Paco Pixmo VSI RefSpatial MMSI VLABench EB-ALFRED EB-Habitat Avg
GPT-4o 47.0 41.8 29.1 29.5 16.2 10.8 42.5 8.8 30.3 39.3 56.3 59.0 34.2
Claude-3.7-Sonnet 35.5 41.3 25.6 22.2 12.4 7.2 47.0 7.7 30.2 41.7 67.0 65.7 33.6
Gemini-2.5-Pro 55.0 42.9 39.9 62.8 45.5 25.8 43.4 30.3 36.9 34.8 62.7 53.0 44.4
InternVL3-2B (init) 31.5 30.9 5.2 7.1 15.4 1.4 31.5 1.8 25.3 19.4 1.3 12.0 15.2
Qwen2.5VL-3B 35.3 30.3 31.0 41.7 67.4 36.6 27.9 24.9 26.5 31.3 6.7 19.7 31.6
RoboBrain2.0-3B 37.3 41.8 64.2 46.0 67.6 36.9 28.8 46.5 26.8 18.1 0.0 10.0 35.3
Vlaser-2B 35.8 38.3 74.0 57.8 72.5 44.6 57.5 43.0 23.6 23.1 42.3 30.7 45.3
InternVL3-8B (init) 35.3 40.0 10.0 14.2 21.1 5.7 42.1 5.6 25.7 24.7 19.0 23.7 22.3
Qwen2.5VL-7B 39.3 29.7 31.1 56.3 68.0 43.5 38.2 32.1 25.9 36.4 10.0 18.3 35.7
Embodied-R1-7B 38.3 37.1 69.5 51.2 69.9 39.2 38.6 31.1 28.1 35.5 10.0 19.0 38.9
RoboBrain2.0-7B 42.0 33.2 63.6 49.5 73.1 37.8 36.1 32.5 26.5 6.6 14.0 29.3 37.0
Vlaser-8B 41.0 53.4 69.5 60.3 68.3 40.5 60.3 59.2 27.2 45.6 50.0 40.0 51.3

Vlaser-2B avg 45.3 beats GPT-4o (34.2) and Claude-3.7-Sonnet (33.6); Vlaser-8B avg 51.3 beats Gemini-2.5-Pro (44.4). +30 pt jump from InternVL3-2B init (15.2 → 45.3) and +29 pt from InternVL3-8B (22.3 → 51.3) demonstrates Vlaser-6M's effectiveness.

SimplerEnv WidowX (Bridge), 4 tasks (Table 2)

Model Carrot Eggplant Spoon Stack Cube Avg
RT-1-X (35M) 4.2 0 0 0 1.1
Octo-Base (93M) 8.3 43.1 12.5 31.9 16.0
OpenVLA (7B) 0 4.1 0 0 1.0
RoboVLM (2B) 25.0 58.3 29.2 12.5 31.3
SpatialVLA (4B) 25.0 100.0 16.7 62.5 42.7
π₀ (3B) 55.8 79.2 63.3 21.3 54.9
InternVL3-2B (no Vlaser data) 42.9 57.1 55.8 11.3 41.8
Vlaser-OOD (2B, OOD reasoning only) 60.8 35.4 56.7 20.0 43.2
Vlaser-QA (2B, + in-domain Bridge QA) 55.8 83.3 77.9 33.3 62.6
Vlaser-Spatial (2B) 48.3 81.7 76.7 36.7 60.8
Vlaser-Grounding (2B) 47.5 80.8 80.0 39.6 62.0
Vlaser-All (2B) 52.5 87.9 76.6 43.3 65.1

SOTA on WidowX (Vlaser-All 65.1 vs π₀'s 54.9 with similar param count). Crucially, the Vlaser-OOD (2B) row trained on out-of-domain Vlaser-6M reasoning data alone is only marginally better than InternVL3-2B (43.2 vs 41.8) — adding in-domain Bridge QA/Spatial/Grounding subsets does the heavy lifting (each ≈ +17-21 pts; Vlaser-QA 62.6, combining all three reaches 65.1).

SimplerEnv Google Robot, 3 tasks (Visual Matching / Variant Aggregation)

Model VM-Pick VM-Move VM-Drawer VM-Avg VA-Pick VA-Move VA-Drawer VA-Avg
TraceVLA (7B) 28.0 53.7 57.0 42.0 60.0 56.4 31.0 45.0
RT-1-X (35M) 56.7 31.7 59.7 53.4 49.0 32.3 29.4 39.6
Octo-Base (93M) 17.0 4.2 22.7 16.8 0.6 3.1 1.1 1.1
OpenVLA (7B) 16.3 46.2 35.6 27.7 54.5 47.7 17.7 39.8
RoboVLM (2B) 77.3 61.7 43.5 63.4 75.6 60.0 10.6 51.3
Emma-X (7B) 2.3 3.3 18.3 8.0 5.3 7.3 20.5 11.0
Magma (8B) 56.0 65.4 83.7 68.4 53.4 65.7 68.8 62.6
GR00T N1.5 (2.1B) 69.3 68.7 35.8 52.4 46.7 62.9 17.5 43.7
π₀ (3B) 72.7 65.3 38.3 58.3 75.2 63.7 25.6 54.8
InternVL3-2B 94.3 78.8 19.0 64.0 80.4 72.7 11.1 54.7
Vlaser (2B) 85.0 76.3 44.9 68.7 74.4 69.2 10.3 51.3
Vlaser-QA (2B) 90.0 84.2 44.4 72.9 78.2 78.2 13.0 56.4
Vlaser-Spatial (2B) 83.0 77.9 56.0 72.3 77.7 73.2 13.2 54.7
Vlaser-Grounding (2B) 83.3 83.3 54.2 73.6 81.2 76.8 17.0 58.3
Vlaser-All (2B) 91.0 85.4 52.1 76.2 80.5 77.7 18.8 59.0

Best Vlaser variant (Vlaser-All, all three in-domain streams combined) reaches VM avg 76.2 and VA avg 59.0 — beating π₀ on VM by +17.9 pts (76.2 vs 58.3). Magma (8B) still leads VA (62.6) via a much stronger Drawer task; Vlaser leads on Pick/Move-Near.

RoboTwin (Aloha-AgileX bimanual, 12 tasks, Table 4)

Model Avg
RDT-1B (re-implemented 30k steps) 36.8%
InternVL3-2B (no Vlaser data) 55.8%
Vlaser-OOD (2B) 54.5%
Vlaser-QA (2B) 60.7%
Vlaser-Spatial (2B) 61.2%
Vlaser-Grounding (2B) 60.7%
Vlaser-All (2B) 67.5%

Same pattern: Vlaser-OOD ≈ InternVL3 baseline; in-domain data drives gains. Best per-task: Click bell 92%, Shake bottle 96%, Place mouse pad 92%.

Ablation Studies

Action prediction hyperparameters (Table 5, WidowX avg)

Model P (predict len) H (execute len) Sample steps δ⁻¹ Avg
InternVL-2B 4 4 10 41.8
InternVL-2B 4 2 10 21.2
InternVL-2B 2 2 10 28.7
InternVL-2B 4 4 20 38.2
Vlaser-OOD (2B) 4 4 10 43.2
Vlaser-QA (2B) 4 4 10 62.6
Vlaser-QA (2B) 4 4 20 63.3

Default (P=4, H=4, 10 steps) is best; doubling sample steps gives marginal gain. Lower H breaks the InternVL baseline harder than it breaks Vlaser — suggesting Vlaser's enhanced state perception is more robust to short execution chunks.

Data-stream contribution (across Tables 2, 3, 4)

The clean ablation split Vlaser-6M into OOD reasoning only (Vlaser-OOD), +QA, +Spatial, +Grounding, and All (the three in-domain types combined). All three in-domain types help individually (≈ +17-21 pts on WidowX vs the 43.2 OOD baseline: QA 62.6, Spatial 60.8, Grounding 62.0), and combining all three pushes WidowX to 65.1% and RoboTwin to 67.5% — diminishing-returns but additive.

Limitations

  1. No real-robot experiments (the single limitation explicitly stated by the authors). Only simulation (SimplerEnv, RoboTwin) so far; real-world validation deferred to future work.
  2. (Reviewer-inferred) Domain gap analysis is descriptive, not prescriptive. Authors observe internet-scale embodied reasoning data ≠ closed-loop control success but don't yet propose a complete fix; in the discussion they argue future work needs alignment techniques to bridge robot-viewpoint vs internet-data viewpoint.
  3. (Reviewer-inferred) Narrow hyperparameter exploration — only 4 (P,H,δ) combinations × 3 model variants tested.

Significance & Positioning

Vs OpenVLA / π₀ / SpatialVLA / RoboVLM (action policy peers). OpenVLA effectively scores 1% on WidowX SimplerEnv (very poor). Vlaser-All-2B at 65.1% beats π₀-3B (54.9%) and SpatialVLA-4B (42.7%) at smaller scale. On Google Robot VM, Vlaser-All beats π₀ by +17.9 pts. Cleanest claim: with the right embodied-reasoning init, a 2B model outperforms 3-8B competitors.

Vs MolmoAct / Embodied-R1 / RoboBrain2.0 / Cosmos-Reason1 (embodied-VLM peers). These work on the upstream VLM but most don't ship a downstream VLA. Vlaser is the first to systematically study upstream→downstream transfer and ship both. Beats Embodied-R1-7B by +12.4 pt avg reasoning, beats RoboBrain2.0-7B by +14.3 pt avg.

Vs GR00T-N1.5 (humanoid foundation). Similar size (Vlaser-2B vs GR00T 2.1B), similar pretraining philosophy, but Vlaser focuses on bimanual + single-arm manipulation in 2 simulators while GR00T targets humanoid full-body. Vlaser beats GR00T on WidowX-style benchmarks; GR00T's strength is humanoid-specific data scale.

Empirical contribution to community knowledge. The "in-domain QA helps, out-of-domain reasoning doesn't" finding contradicts the prevailing intuition that better VLM benchmark scores → better VLA. This shifts the design lever for future VLAs from "scale embodied-reasoning data" to "align VLM observation distribution with the target robot's viewpoint". The argument is that the visual-domain mismatch (internet third-person → robot front-view simulation) bottlenecks transfer more than reasoning capability does.

Vs Embodied-R1 / OneTwoVLA (reasoning-augmented VLAs). Those add reasoning into the action policy itself. Vlaser argues for front-loading reasoning into the VLM stage then doing standard flow-matching action expertise — keeps inference fast and policy clean.

Open release of weights, data engine, and the full Vlaser-6M dataset is non-trivial — the dataset itself becomes a community resource on par with Vlaser-6M's competitor curation efforts (Cosmos-Reason1, RoboBrain2.0, EmbodiedOneVision).

Links

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️