ICLR 2026 InstructVLA - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 · arXiv: 2507.17520 · GitHub: InternRobotics/InstructVLA Affiliations: Shanghai AI Lab, USTC, Zhejiang University · Base VLM: Eagle2-2B Category: Training Approach Trend tag: Trend 2
flowchart TB
Stage1[Stage 1: Action pretraining<br/>heterogeneous manipulation data<br/>actions + rule-based language motion] --> Stage2[Stage 2: VLA-IT<br/>freeze action expert,<br/>add language LoRA + scale head]
Stage2 --> MoE{MoE-adaptation<br/>LoRA experts in LLM}
MoE -- scale head predicts gating λᵢ --> Blend[Adaptively blend expert outputs<br/>reasoning ↔ action]
Data[VLA-IT 650K dataset<br/>instructions, captions, QA pairs] --> Stage2
Like Actions as Language: existing VLA models sacrifice multimodal reasoning for task-specific manipulation and suffer catastrophic forgetting of pretrained vision-language capability. InstructVLA aims to preserve the flexible reasoning of the base VLM while delivering leading manipulation performance, using embodied reasoning to help action.
- Base VLM: Eagle2-2B backbone.
- Two-stage pipeline (order matters): Stage 1 = action pretraining on heterogeneous manipulation data, jointly predicting actions (flow-matching objective) and rule-based annotated language motion (LM loss). Stage 2 = Vision-Language-Action Instruction Tuning (VLA-IT), which freezes the action expert and adds a new language LoRA adapter plus a scale head of the MoE-adaptation.
- MoE-adaptation (not a dual system): LoRA modules act as experts inside the LLM backbone. A scale head predicts gating coefficients λᵢ per expert by classifying the hidden state, adaptively blending their outputs to switch between reasoning and action generation. It does not route separate token types to separate experts; a single VLM emits both text and latent actions.
- VLA-IT 650K: 650K human-robot interactions annotated with diverse instructions, scene captions, and grounded QA pairs, trained jointly with standard VLM corpora.
- In-domain SimplerEnv: +33% over SpatialVLA.
- SimplerEnv-Instruct (newly introduced 80-task benchmark requiring closed-loop control + high-level instruction understanding): outperforms a fine-tuned OpenVLA by 96% and an action expert aided by GPT-4o by 29%.
- Surpasses baseline VLMs on multimodal tasks and shows inference-time scaling — textual reasoning boosts manipulation in sim and real.
Contributes both a method (MoE-adaptation that preserves VLM reasoning during action learning) and a dataset (VLA-IT 650K) reusable by other groups. Demonstrates that embodied reasoning can be co-trained with action generation without sacrificing pretrained multimodal capability, enabling steerable instruction following.
- Actions as Language (the data-relabeling counterpart)
- Embodied-R1 (RL-based reasoning)
← Back to ICLR-2026