ICLR 2026 Actions as Language - Heungwoo/research GitHub Wiki

Actions as Language — Fine-Tuning VLMs into VLAs Without Catastrophic Forgetting

Venue: ICLR 2026 (arXiv:2509.22195, submitted 2025-09-26) Method name: VLM2VLA Authors: Asher J. Hancock, Xindi Wu, Lihan Zha, Olga Russakovsky, Anirudha Majumdar (Princeton University) Base VLM: Gemma-3-12B-IT (also tested with a Prismatic VLM) Category: Training Approach Trend tag: Trend 2

Approach diagram

flowchart LR
  D[Existing teleop dataset] --> Re[Relabel with subtask text<br/>+ intermediate motion plans]
  Re --> Stream[Token stream:<br/>language ↔ actions ↔ language ↔ ...]
  Stream --> LM[Standard language modeling loss<br/>actions are 'just another modality']
  LM --> M[Fine-tuned VLA<br/>preserves VQA + reasoning]
Loading

Problem

Fine-tuning a VLM into a VLA typically causes catastrophic forgetting: the resulting model's VQA and language capabilities degrade. The paper attributes this to a distribution mismatch between the VLM's internet-scale pretraining corpus and the robotics fine-tuning data — full-parameter updates on action prediction overwrite general-purpose world knowledge in the shared backbone. Users were forced to choose between strong language reasoning and strong action prediction.

Method

Resolve the mismatch at the data level: re-represent low-level actions as natural language descriptions so that robot data looks like ordinary VLM (VQA-style) data. Each trajectory is decomposed into sub-trajectories, and action prediction is framed as a three-stage hierarchical VQA reasoning process (subtask → intermediate motion plan → action), all expressed in language. Because the fine-tuning data now matches the VLM's pretraining distribution, the model can be adapted using LoRA alone, minimally perturbing the backbone and averting catastrophic forgetting. No new action tokenizer or architectural head is added.

Results

Validated with over 800 real-world manipulation experiments plus VQA studies across 12 multimodal benchmarks (MMMU, MMStar, MME, OCRBench, MMB-en/cn, TextVQA, DocVQA, InfoVQA, AI2D, ChartQA, RealWorldQA). VLM2VLA preserves base-VLM understanding (e.g., 42.7 MMMU, 48.0 MMStar, comparable to Gemma-3-12B-IT) while remaining competitive at manipulation. Preserved language reasoning enables zero-shot generalization to novel semantic tasks and multilingual instruction following (Hindi, Mandarin demonstrated), outperforming OpenVLA and ECoT on out-of-distribution tasks.

Significance

Reframes VLA fine-tuning as a data-alignment / instruction-tuning problem rather than a modality-injection problem: by making robot data linguistically indistinguishable from VLM pretraining data, a lightweight LoRA adapter suffices and the backbone is barely touched. No architectural change, just relabeling actions into language. With InstructVLA this paper supports the view that catastrophic forgetting in VLA fine-tuning is largely addressable.

Links

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️