CoRL 2025 pi05 - Heungwoo/research GitHub Wiki

ฯ€0.5 โ€” A VLA Model with Open-World Generalization

Venue: CoRL 2025 (Oral) ยท Author: Physical Intelligence ยท arXiv: 2504.16054 Category: VLA Architecture (Baseline) Trend tag: Hierarchy wins in the open world

Approach diagram

flowchart LR
  V[Vision] --> VLM[VLM Backbone<br/>web-pretrained]
  L[Task: 'clean the kitchen'] --> VLM
  VLM --> HL[High-level<br/>subtask predictor<br/>'open the drawer']
  VLM --> AE[Flow-Matching<br/>Action Expert<br/>continuous chunks]
  HL --> AE
  noise --> AE
  subgraph Co-training sources
    WD[Web VQA + captions]
    DET[Object detection]
    SUB[Subtask semantic prediction]
    TELE[Multi-robot teleop]
  end
  WD & DET & SUB & TELE -. mixed batches .-> VLM
Loading

Problem

ฯ€0 was a cross-embodiment flow-matching VLA that worked on in-distribution tasks. But dropping it into an entirely new home โ€” cleaning an unfamiliar kitchen or bedroom โ€” broke generalization: the model couldn't parse novel scene layouts and couldn't decompose long-horizon housework into executable subtasks.

Method

Two additions on top of ฯ€0:

  1. Hierarchical split. A high-level head predicts the next semantic subtask (โ„“ฬ‚, e.g., "open the fridge door") via autoregressive token decoding; a low-level ~300M-param flow-matching expert generates continuous action chunks (50-step / 1-second) conditioned on both the overall task โ„“ and the subtask โ„“ฬ‚.
  2. Heterogeneous co-training. Mixed batches include multi-robot teleop + web VQA/captions + object-detection tasks + subtask-prediction tasks. Pretraining happens with FAST action tokens (discrete, for gradient stability), then flow-matching for the action expert.

Data scale: ~400 hrs mobile-manipulator data ร— ~100 homes, plus cross-embodiment lab data and OXE; 97.6% of pretraining examples are from non-MM sources.

Results

First end-to-end learning-enabled system to do long-horizon, dexterous manipulation in entirely unseen homes โ€” cleaning kitchens, tidying bedrooms. In out-of-distribution homes ฯ€0.5 reaches ~94% task success and ~94% language-following, approaching baselines trained directly on the test environments, with quantitatively significant jumps over ฯ€0 on open-world generalization suites; task-specific post-training still often needed for polish (relaxed later by ฯ€0.6 and ฯ€0.7).

Significance

ฯ€0.5 is the ancestor of the entire current ฯ€ series. Every subsequent release builds on it:

  • ฯ€0.6 upgrades backbone to Gemma3-4B + Knowledge Insulation.
  • ฯ€*0.6 + RECAP adds advantage-conditioned RL for flow-matching.
  • ฯ€0.7 adds MEM video history, subgoal-image world-model conditioning, and metadata prompting.

See ฯ€ series evolution for the full side-by-side. ฯ€0.5 also establishes the "hierarchical generalist" template that OneTwoVLA, Long-VLA, and ICLR 2026's reasoning-augmented VLAs follow.

Links

Related pages

โ† Back to CoRL-2025

โš ๏ธ **GitHub.com Fallback** โš ๏ธ