CoRL 2025 pi05 - Heungwoo/research GitHub Wiki
Venue: CoRL 2025 (Oral) ยท Author: Physical Intelligence ยท arXiv: 2504.16054 Category: VLA Architecture (Baseline) Trend tag: Hierarchy wins in the open world
flowchart LR
V[Vision] --> VLM[VLM Backbone<br/>web-pretrained]
L[Task: 'clean the kitchen'] --> VLM
VLM --> HL[High-level<br/>subtask predictor<br/>'open the drawer']
VLM --> AE[Flow-Matching<br/>Action Expert<br/>continuous chunks]
HL --> AE
noise --> AE
subgraph Co-training sources
WD[Web VQA + captions]
DET[Object detection]
SUB[Subtask semantic prediction]
TELE[Multi-robot teleop]
end
WD & DET & SUB & TELE -. mixed batches .-> VLM
ฯ0 was a cross-embodiment flow-matching VLA that worked on in-distribution tasks. But dropping it into an entirely new home โ cleaning an unfamiliar kitchen or bedroom โ broke generalization: the model couldn't parse novel scene layouts and couldn't decompose long-horizon housework into executable subtasks.
Two additions on top of ฯ0:
- Hierarchical split. A high-level head predicts the next semantic subtask (โฬ, e.g., "open the fridge door") via autoregressive token decoding; a low-level ~300M-param flow-matching expert generates continuous action chunks (50-step / 1-second) conditioned on both the overall task โ and the subtask โฬ.
- Heterogeneous co-training. Mixed batches include multi-robot teleop + web VQA/captions + object-detection tasks + subtask-prediction tasks. Pretraining happens with FAST action tokens (discrete, for gradient stability), then flow-matching for the action expert.
Data scale: ~400 hrs mobile-manipulator data ร ~100 homes, plus cross-embodiment lab data and OXE; 97.6% of pretraining examples are from non-MM sources.
First end-to-end learning-enabled system to do long-horizon, dexterous manipulation in entirely unseen homes โ cleaning kitchens, tidying bedrooms. In out-of-distribution homes ฯ0.5 reaches ~94% task success and ~94% language-following, approaching baselines trained directly on the test environments, with quantitatively significant jumps over ฯ0 on open-world generalization suites; task-specific post-training still often needed for polish (relaxed later by ฯ0.6 and ฯ0.7).
ฯ0.5 is the ancestor of the entire current ฯ series. Every subsequent release builds on it:
- ฯ0.6 upgrades backbone to Gemma3-4B + Knowledge Insulation.
- ฯ*0.6 + RECAP adds advantage-conditioned RL for flow-matching.
- ฯ0.7 adds MEM video history, subgoal-image world-model conditioning, and metadata prompting.
See ฯ series evolution for the full side-by-side. ฯ0.5 also establishes the "hierarchical generalist" template that OneTwoVLA, Long-VLA, and ICLR 2026's reasoning-augmented VLAs follow.
- arXiv: https://arxiv.org/abs/2504.16054
- Project PDF: https://www.pi.website/download/pi05.pdf
- Blog: https://www.pi.website/blog/pi05
- OpenPi: https://github.com/Physical-Intelligence/openpi
โ Back to CoRL-2025