ICML 2026 Learning to Move Before Learning to Do - Heungwoo/research GitHub Wiki

Learning to Move Before Learning to Do — task-agnostic pretraining for VLAs

Venue: ICML 2026 (Poster) Category: VLA Architecture

Problem

Vision-Language-Action (VLA) models are "bottlenecked by the scarcity of expert demonstrations—expensive triplets of observations, language instructions, and actions." Collecting aligned (observation, instruction, action) triplets is costly and slow, capping how far standard behavior cloning can scale. The paper's core hypothesis is that learning how to move can be decoupled from learning what to do, and that the former needs no task labels at all.

Method

The authors propose Task-Agnostic Pretraining (TAP), a two-stage framework:

  • Stage 1 — task-agnostic pretraining. Pre-train on abundant, cheap task-agnostic data—discarded off-task trajectories or autonomous robot play—using an Inverse Dynamics objective that predicts the action between two consecutive observations. This self-supervised phase requires no human annotation and "instills physical affordances—grasping, contact dynamics, end-effector control."
  • Stage 2 — lightweight language alignment. A lightweight second stage then aligns the learned physical priors with language instructions using only minimal expert data, converting the task-agnostic motion prior into an instruction-following policy.

The key insight is that inverse dynamics turns unlabeled, low-value play/off-task data into a strong, transferable physical representation, so the scarce labeled triplets are spent only on language grounding rather than on relearning motor control.

flowchart LR
    A[Cheap task-agnostic data:\noff-task + robot play] --> B[Stage 1: Inverse Dynamics\npredict action from o_t, o_t+1]
    B --> C[Physical priors:\ngrasping, contact, EEF control]
    C --> D[Stage 2: lightweight\nlanguage alignment\nminimal expert data]
    D --> E[Instruction-following VLA]

Results

On the SIMPLER benchmark, TAP "matches models trained on 1M+ expert trajectories while using orders of magnitude less labeled data, achieving a 10% absolute gain over standard behavior cloning." In real-world WidowX experiments it surpasses internet-scale baselines under visual distribution shifts, with the headline result of 25% vs. 0% under camera perturbations.

Significance

TAP reframes the VLA data bottleneck: instead of chasing more labeled triplets, it mines the cheap, normally-discarded off-task and play data that robots already generate, and shows inverse-dynamics pretraining yields robust, transferable motor priors. The strong robustness under camera perturbation (25% vs. 0%) suggests task-agnostic physical pretraining gives better distribution-shift generalization than internet-scale supervised baselines—an attractive recipe for embodied AI where labeled data is the scarce resource.

Links

← Back to ICML-2026