ICLR 2026 Actions as Language - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 (arXiv:2509.22195, submitted 2025-09-26) Method name: VLM2VLA Authors: Asher J. Hancock, Xindi Wu, Lihan Zha, Olga Russakovsky, Anirudha Majumdar (Princeton University) Base VLM: Gemma-3-12B-IT (also tested with a Prismatic VLM) Category: Training Approach Trend tag: Trend 2
flowchart LR
D[Existing teleop dataset] --> Re[Relabel with subtask text<br/>+ intermediate motion plans]
Re --> Stream[Token stream:<br/>language ↔ actions ↔ language ↔ ...]
Stream --> LM[Standard language modeling loss<br/>actions are 'just another modality']
LM --> M[Fine-tuned VLA<br/>preserves VQA + reasoning]
Fine-tuning a VLM into a VLA typically causes catastrophic forgetting: the resulting model's VQA and language capabilities degrade. The paper attributes this to a distribution mismatch between the VLM's internet-scale pretraining corpus and the robotics fine-tuning data — full-parameter updates on action prediction overwrite general-purpose world knowledge in the shared backbone. Users were forced to choose between strong language reasoning and strong action prediction.
Resolve the mismatch at the data level: re-represent low-level actions as natural language descriptions so that robot data looks like ordinary VLM (VQA-style) data. Each trajectory is decomposed into sub-trajectories, and action prediction is framed as a three-stage hierarchical VQA reasoning process (subtask → intermediate motion plan → action), all expressed in language. Because the fine-tuning data now matches the VLM's pretraining distribution, the model can be adapted using LoRA alone, minimally perturbing the backbone and averting catastrophic forgetting. No new action tokenizer or architectural head is added.
Validated with over 800 real-world manipulation experiments plus VQA studies across 12 multimodal benchmarks (MMMU, MMStar, MME, OCRBench, MMB-en/cn, TextVQA, DocVQA, InfoVQA, AI2D, ChartQA, RealWorldQA). VLM2VLA preserves base-VLM understanding (e.g., 42.7 MMMU, 48.0 MMStar, comparable to Gemma-3-12B-IT) while remaining competitive at manipulation. Preserved language reasoning enables zero-shot generalization to novel semantic tasks and multilingual instruction following (Hindi, Mandarin demonstrated), outperforming OpenVLA and ECoT on out-of-distribution tasks.
Reframes VLA fine-tuning as a data-alignment / instruction-tuning problem rather than a modality-injection problem: by making robot data linguistically indistinguishable from VLM pretraining data, a lightweight LoRA adapter suffices and the backbone is barely touched. No architectural change, just relabeling actions into language. With InstructVLA this paper supports the view that catastrophic forgetting in VLA fine-tuning is largely addressable.
- arXiv: https://arxiv.org/abs/2509.22195
- Project page: https://vlm2vla.github.io/
- OpenReview: https://openreview.net/forum?id=sFO9d6XSlf
- InstructVLA (architectural counterpart)
← Back to ICLR-2026