ICLR 2026 Vlaser - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 Authors: Shanghai AI Lab + USTC + SJTU + Zhejiang + Nanjing + Fudan + Tsinghua + NUS + Northeastern + Shenzhen U. (Ganlin Yang, Tianyi Zhang, Haoran Hao, Weiyun Wang, Yibin Liu, … Yao Mu, Zhi Hou) Category: VLA Architecture — Reasoning-augmented · VLM-init study Trend tag: Embodied reasoning · VLM→VLA initialization · domain-shift mitigation
flowchart LR
Net[InternVL3 backbone<br/>2B / 8B] --> SFT[Stage 1: VLM SFT<br/>Vlaser-6M dataset]
D6M[Vlaser-6M:<br/>1.5M grounding · 1.2M RoboVQA<br/>0.5M spatial · 0.4M planning<br/>2M in-domain SimplerEnv + RoboTwin2.0] --> SFT
SFT --> VLM[Vlaser VLM<br/>SOTA on 12 embodied benchmarks]
VLM --> AE[Action Expert<br/>flow-matching, π0-style<br/>shared self-attention<br/>non-causal VLA stream]
Obs[image + lang + robot state q_t] --> AE
AE --> A[Action chunk A_t<br/>H=4, 10 inference steps Euler]
VLAs are typically built by fine-tuning generic web-pretrained VLMs (Kim '24 OpenVLA, Black '24 π₀, Qu '25 SpatialVLA), but few studies examine how the VLM initialization choice affects downstream VLA. The hidden assumption is that reasoning-stronger VLMs make stronger VLAs — Vlaser tests it. They find a counter-intuitive result: out-of-domain embodied reasoning data does NOT translate cleanly to closed-loop manipulation; what helps is in-domain data from the same simulator/embodiment.
- VLM backbone: InternVL3-2B (InternViT + Qwen2.5-1.5B) and InternVL3-8B (InternViT + Qwen2.5-7B). Two model sizes for compute-vs-quality trade-offs.
- Action expert: π₀-style mixture-of-experts head — original VLM weights handle image/text streams; a separate weight set handles robot state + noised actions. Self-attention is shared between language and action experts; non-causal attention for the VLA stream.
- Robot state q_t encoded as a state token; noised actions A_t^τ as action tokens; both projected via linear layer into the same embedding space as image/text features.
Six categories, 6M total samples:
| Category | Size | Sources |
|---|---|---|
| Embodied grounding (bbox + center points, normalized [0,1000]) | 1.5M (+0.3M from SA-1B masks) | RoboPoint, ShareRobot, Pixmo-Points, Paco-LaVIS, RefSpatial |
| RoboVQA (general state perception) | 1.2M | RoboVQA, Robo2VLM, RoboPoint, RefSpatial, OWMM-Agent |
| Spatial intelligence | 0.5M (+0.1M manually annotated) | SPAR, SpaceR-151k, VILASR; ScanNet, ScanNet++, CA-1M, ARKitScenes |
| Planning | 0.4M | Alpaca-15k, MuEP, WAP, LLaRP/Habitat, EgoPlan-IT, EgoCOT |
| In-domain VLA data (SimplerEnv Google Robot + WidowX, and RoboTwin2.0 Aloha-AgileX) | 2M | Synthesized QA on same observation space as eval; LLM-as-judge (Qwen2.5VL-32B) filtering drops ~10% lowest-scored |
| Total | ~6M |
The "in-domain 2M" subset is the key empirical lever — synthesized QA annotations directly on robot observation footage from the SimplerEnv simulator (WidowX + Google Robot) and RoboTwin2.0 (Aloha-AgileX), in the reasoning categories above.
Stage 1 — Vision-Language Pretraining (SFT on Vlaser-6M):
- Auto-regressive language-modeling loss L_lm = -log p(t_N | F_v(x; θ_v), F_t(y), t_{0:N-1}; Θ).
- Initialized from InternVL3-2B / InternVL3-8B; full-parameter SFT.
Stage 2 — VLA fine-tuning (flow matching):
- Action chunk A_t = [a_t, …, a_{t+H-1}] with horizon H = 4 (Action Chunk length = 4, Table 7).
- Noisy action chunk A_t^τ = τA_t + (1-τ)ε, ε ~ N(0, I).
- Loss L_vla = E_{p(A_t|o_t)} ‖v_θ(A_t^τ, o_t) - u(A_t^τ | A_t)‖² where u = ε - A_t.
- Inference: forward Euler numerical integration, 10 inference steps; observation history length = 1 (single-frame). At eval only 2 of the 4 predicted actions are executed open-loop before re-planning (Execute Action length = 2, Table 7 / eval appendix) — note the Table 5 ablation separately reports a default "execute length H=4", an internal terminology clash in the paper.
- Eval on SimplerEnv (WidowX Bridge + Google Robot) and RoboTwin 2.0 (Aloha-AgileX bimanual, Table 4). No real-robot experiments.
Training hyperparameters (Appendix Tables 6–7): VLM SFT — AdamW, peak LR 2e-5, cosine decay, 1 epoch / 5,000 steps, global batch 128, 150 warm-up steps, bf16, LLM seq len 16,384, dynamic resolution (patch 448, max 12 patches), all params trainable. VLA fine-tuning — AdamW, VLM & action-expert peak LR 5e-5, global batch 1,024, 10 epochs, seq len 384, all params trainable; WidowX checkpoint at 45,390 iters, Google Robot at 36,970 iters. (GPU hours not reported.)
| ERQA | EgoPlan2 | Where2place | Pointarena | Paco | Pixmo | VSI | RefSpatial | MMSI | VLABench | EB-ALFRED | EB-Habitat | Avg | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| GPT-4o | 47.0 | 41.8 | 29.1 | 29.5 | 16.2 | 10.8 | 42.5 | 8.8 | 30.3 | 39.3 | 56.3 | 59.0 | 34.2 |
| Claude-3.7-Sonnet | 35.5 | 41.3 | 25.6 | 22.2 | 12.4 | 7.2 | 47.0 | 7.7 | 30.2 | 41.7 | 67.0 | 65.7 | 33.6 |
| Gemini-2.5-Pro | 55.0 | 42.9 | 39.9 | 62.8 | 45.5 | 25.8 | 43.4 | 30.3 | 36.9 | 34.8 | 62.7 | 53.0 | 44.4 |
| InternVL3-2B (init) | 31.5 | 30.9 | 5.2 | 7.1 | 15.4 | 1.4 | 31.5 | 1.8 | 25.3 | 19.4 | 1.3 | 12.0 | 15.2 |
| Qwen2.5VL-3B | 35.3 | 30.3 | 31.0 | 41.7 | 67.4 | 36.6 | 27.9 | 24.9 | 26.5 | 31.3 | 6.7 | 19.7 | 31.6 |
| RoboBrain2.0-3B | 37.3 | 41.8 | 64.2 | 46.0 | 67.6 | 36.9 | 28.8 | 46.5 | 26.8 | 18.1 | 0.0 | 10.0 | 35.3 |
| Vlaser-2B | 35.8 | 38.3 | 74.0 | 57.8 | 72.5 | 44.6 | 57.5 | 43.0 | 23.6 | 23.1 | 42.3 | 30.7 | 45.3 |
| InternVL3-8B (init) | 35.3 | 40.0 | 10.0 | 14.2 | 21.1 | 5.7 | 42.1 | 5.6 | 25.7 | 24.7 | 19.0 | 23.7 | 22.3 |
| Qwen2.5VL-7B | 39.3 | 29.7 | 31.1 | 56.3 | 68.0 | 43.5 | 38.2 | 32.1 | 25.9 | 36.4 | 10.0 | 18.3 | 35.7 |
| Embodied-R1-7B | 38.3 | 37.1 | 69.5 | 51.2 | 69.9 | 39.2 | 38.6 | 31.1 | 28.1 | 35.5 | 10.0 | 19.0 | 38.9 |
| RoboBrain2.0-7B | 42.0 | 33.2 | 63.6 | 49.5 | 73.1 | 37.8 | 36.1 | 32.5 | 26.5 | 6.6 | 14.0 | 29.3 | 37.0 |
| Vlaser-8B | 41.0 | 53.4 | 69.5 | 60.3 | 68.3 | 40.5 | 60.3 | 59.2 | 27.2 | 45.6 | 50.0 | 40.0 | 51.3 |
Vlaser-2B avg 45.3 beats GPT-4o (34.2) and Claude-3.7-Sonnet (33.6); Vlaser-8B avg 51.3 beats Gemini-2.5-Pro (44.4). +30 pt jump from InternVL3-2B init (15.2 → 45.3) and +29 pt from InternVL3-8B (22.3 → 51.3) demonstrates Vlaser-6M's effectiveness.
| Model | Carrot | Eggplant | Spoon | Stack Cube | Avg |
|---|---|---|---|---|---|
| RT-1-X (35M) | 4.2 | 0 | 0 | 0 | 1.1 |
| Octo-Base (93M) | 8.3 | 43.1 | 12.5 | 31.9 | 16.0 |
| OpenVLA (7B) | 0 | 4.1 | 0 | 0 | 1.0 |
| RoboVLM (2B) | 25.0 | 58.3 | 29.2 | 12.5 | 31.3 |
| SpatialVLA (4B) | 25.0 | 100.0 | 16.7 | 62.5 | 42.7 |
| π₀ (3B) | 55.8 | 79.2 | 63.3 | 21.3 | 54.9 |
| InternVL3-2B (no Vlaser data) | 42.9 | 57.1 | 55.8 | 11.3 | 41.8 |
| Vlaser-OOD (2B, OOD reasoning only) | 60.8 | 35.4 | 56.7 | 20.0 | 43.2 |
| Vlaser-QA (2B, + in-domain Bridge QA) | 55.8 | 83.3 | 77.9 | 33.3 | 62.6 |
| Vlaser-Spatial (2B) | 48.3 | 81.7 | 76.7 | 36.7 | 60.8 |
| Vlaser-Grounding (2B) | 47.5 | 80.8 | 80.0 | 39.6 | 62.0 |
| Vlaser-All (2B) | 52.5 | 87.9 | 76.6 | 43.3 | 65.1 |
SOTA on WidowX (Vlaser-All 65.1 vs π₀'s 54.9 with similar param count). Crucially, the Vlaser-OOD (2B) row trained on out-of-domain Vlaser-6M reasoning data alone is only marginally better than InternVL3-2B (43.2 vs 41.8) — adding in-domain Bridge QA/Spatial/Grounding subsets does the heavy lifting (each ≈ +17-21 pts; Vlaser-QA 62.6, combining all three reaches 65.1).
| Model | VM-Pick | VM-Move | VM-Drawer | VM-Avg | VA-Pick | VA-Move | VA-Drawer | VA-Avg |
|---|---|---|---|---|---|---|---|---|
| TraceVLA (7B) | 28.0 | 53.7 | 57.0 | 42.0 | 60.0 | 56.4 | 31.0 | 45.0 |
| RT-1-X (35M) | 56.7 | 31.7 | 59.7 | 53.4 | 49.0 | 32.3 | 29.4 | 39.6 |
| Octo-Base (93M) | 17.0 | 4.2 | 22.7 | 16.8 | 0.6 | 3.1 | 1.1 | 1.1 |
| OpenVLA (7B) | 16.3 | 46.2 | 35.6 | 27.7 | 54.5 | 47.7 | 17.7 | 39.8 |
| RoboVLM (2B) | 77.3 | 61.7 | 43.5 | 63.4 | 75.6 | 60.0 | 10.6 | 51.3 |
| Emma-X (7B) | 2.3 | 3.3 | 18.3 | 8.0 | 5.3 | 7.3 | 20.5 | 11.0 |
| Magma (8B) | 56.0 | 65.4 | 83.7 | 68.4 | 53.4 | 65.7 | 68.8 | 62.6 |
| GR00T N1.5 (2.1B) | 69.3 | 68.7 | 35.8 | 52.4 | 46.7 | 62.9 | 17.5 | 43.7 |
| π₀ (3B) | 72.7 | 65.3 | 38.3 | 58.3 | 75.2 | 63.7 | 25.6 | 54.8 |
| InternVL3-2B | 94.3 | 78.8 | 19.0 | 64.0 | 80.4 | 72.7 | 11.1 | 54.7 |
| Vlaser (2B) | 85.0 | 76.3 | 44.9 | 68.7 | 74.4 | 69.2 | 10.3 | 51.3 |
| Vlaser-QA (2B) | 90.0 | 84.2 | 44.4 | 72.9 | 78.2 | 78.2 | 13.0 | 56.4 |
| Vlaser-Spatial (2B) | 83.0 | 77.9 | 56.0 | 72.3 | 77.7 | 73.2 | 13.2 | 54.7 |
| Vlaser-Grounding (2B) | 83.3 | 83.3 | 54.2 | 73.6 | 81.2 | 76.8 | 17.0 | 58.3 |
| Vlaser-All (2B) | 91.0 | 85.4 | 52.1 | 76.2 | 80.5 | 77.7 | 18.8 | 59.0 |
Best Vlaser variant (Vlaser-All, all three in-domain streams combined) reaches VM avg 76.2 and VA avg 59.0 — beating π₀ on VM by +17.9 pts (76.2 vs 58.3). Magma (8B) still leads VA (62.6) via a much stronger Drawer task; Vlaser leads on Pick/Move-Near.
| Model | Avg |
|---|---|
| RDT-1B (re-implemented 30k steps) | 36.8% |
| InternVL3-2B (no Vlaser data) | 55.8% |
| Vlaser-OOD (2B) | 54.5% |
| Vlaser-QA (2B) | 60.7% |
| Vlaser-Spatial (2B) | 61.2% |
| Vlaser-Grounding (2B) | 60.7% |
| Vlaser-All (2B) | 67.5% |
Same pattern: Vlaser-OOD ≈ InternVL3 baseline; in-domain data drives gains. Best per-task: Click bell 92%, Shake bottle 96%, Place mouse pad 92%.
| Model | P (predict len) | H (execute len) | Sample steps δ⁻¹ | Avg |
|---|---|---|---|---|
| InternVL-2B | 4 | 4 | 10 | 41.8 |
| InternVL-2B | 4 | 2 | 10 | 21.2 |
| InternVL-2B | 2 | 2 | 10 | 28.7 |
| InternVL-2B | 4 | 4 | 20 | 38.2 |
| Vlaser-OOD (2B) | 4 | 4 | 10 | 43.2 |
| Vlaser-QA (2B) | 4 | 4 | 10 | 62.6 |
| Vlaser-QA (2B) | 4 | 4 | 20 | 63.3 |
Default (P=4, H=4, 10 steps) is best; doubling sample steps gives marginal gain. Lower H breaks the InternVL baseline harder than it breaks Vlaser — suggesting Vlaser's enhanced state perception is more robust to short execution chunks.
The clean ablation split Vlaser-6M into OOD reasoning only (Vlaser-OOD), +QA, +Spatial, +Grounding, and All (the three in-domain types combined). All three in-domain types help individually (≈ +17-21 pts on WidowX vs the 43.2 OOD baseline: QA 62.6, Spatial 60.8, Grounding 62.0), and combining all three pushes WidowX to 65.1% and RoboTwin to 67.5% — diminishing-returns but additive.
- No real-robot experiments (the single limitation explicitly stated by the authors). Only simulation (SimplerEnv, RoboTwin) so far; real-world validation deferred to future work.
- (Reviewer-inferred) Domain gap analysis is descriptive, not prescriptive. Authors observe internet-scale embodied reasoning data ≠ closed-loop control success but don't yet propose a complete fix; in the discussion they argue future work needs alignment techniques to bridge robot-viewpoint vs internet-data viewpoint.
- (Reviewer-inferred) Narrow hyperparameter exploration — only 4 (P,H,δ) combinations × 3 model variants tested.
Vs OpenVLA / π₀ / SpatialVLA / RoboVLM (action policy peers). OpenVLA effectively scores 1% on WidowX SimplerEnv (very poor). Vlaser-All-2B at 65.1% beats π₀-3B (54.9%) and SpatialVLA-4B (42.7%) at smaller scale. On Google Robot VM, Vlaser-All beats π₀ by +17.9 pts. Cleanest claim: with the right embodied-reasoning init, a 2B model outperforms 3-8B competitors.
Vs MolmoAct / Embodied-R1 / RoboBrain2.0 / Cosmos-Reason1 (embodied-VLM peers). These work on the upstream VLM but most don't ship a downstream VLA. Vlaser is the first to systematically study upstream→downstream transfer and ship both. Beats Embodied-R1-7B by +12.4 pt avg reasoning, beats RoboBrain2.0-7B by +14.3 pt avg.
Vs GR00T-N1.5 (humanoid foundation). Similar size (Vlaser-2B vs GR00T 2.1B), similar pretraining philosophy, but Vlaser focuses on bimanual + single-arm manipulation in 2 simulators while GR00T targets humanoid full-body. Vlaser beats GR00T on WidowX-style benchmarks; GR00T's strength is humanoid-specific data scale.
Empirical contribution to community knowledge. The "in-domain QA helps, out-of-domain reasoning doesn't" finding contradicts the prevailing intuition that better VLM benchmark scores → better VLA. This shifts the design lever for future VLAs from "scale embodied-reasoning data" to "align VLM observation distribution with the target robot's viewpoint". The argument is that the visual-domain mismatch (internet third-person → robot front-view simulation) bottlenecks transfer more than reasoning capability does.
Vs Embodied-R1 / OneTwoVLA (reasoning-augmented VLAs). Those add reasoning into the action policy itself. Vlaser argues for front-loading reasoning into the VLM stage then doing standard flow-matching action expertise — keeps inference fast and policy clean.
Open release of weights, data engine, and the full Vlaser-6M dataset is non-trivial — the dataset itself becomes a community resource on par with Vlaser-6M's competitor curation efforts (Cosmos-Reason1, RoboBrain2.0, EmbodiedOneVision).
- OpenReview: https://openreview.net/forum?id=8xTDnj39Ti
- GitHub: https://github.com/OpenGVLab/Vlaser/
← Back to ICLR-2026