ICRA 2026 Galaxea G0 - Heungwoo/research GitHub Wiki

Galaxea + G0 — Open-World Dataset & Dual-System VLA

(The ICRA 2026 program lists the title as "Galaxy Open-World Dataset…"; the authoritative paper/brand is Galaxea (Galaxea AI / OpenGalaxea), used throughout below.)

Venue: ICRA 2026 · Authors: Tao Jiang, Tianyuan Yuan, Yicheng Liu, Chenhao Lu, Jianning Cui, Xiao Liu, Shuiqi Cheng, Jiyang Gao, Huazhe Xu, Hang Zhao (Galaxea Team / Tsinghua) · arXiv: 2509.00576 Category: Data + dual-system VLA Trend tag: Open-world dataset + System 1/2 VLA

Approach diagram

flowchart LR
  subgraph Data[Galaxea Open-World Dataset]
    EMB[Galaxea R1 Lite<br/>23-DoF mobile bimanual] --> COLL[500 h · 100K trajectories<br/>50 scenes · 11 sites]
    COLL --> ANN[Subtask-level<br/>language annotations]
  end
  ANN --> S2[System-2: G0-VLM<br/>Qwen2.5-VL planner<br/>high-level subtask reasoning]
  IMG[Multi-view RGB + instruction] --> S2
  S2 -->|subtask goal| S1[System-1: G0-VLA<br/>PaliGemma 3B + SigLIP<br/>action expert]
  IMG --> S1
  S1 --> ACT[Action chunk → bimanual control]
Loading

Problem

Most robot-learning datasets mix embodiments or lack fine-grained language, and most VLAs either reason slowly or act reactively but not both. The authors target open-world mobile manipulation in authentic homes, kitchens, retail and office spaces — settings that demand both long-horizon planning and precise, real-time bimanual control. Two gaps: (1) a large, consistent-embodiment dataset with subtask annotations, and (2) an architecture that couples deliberative planning (System 2) with fast execution (System 1) without either bottlenecking the other.

Method

Galaxea Open-World Dataset. 500 hours of high-fidelity real-world data, 100K demonstration trajectories spanning 150 task categories, collected across 50 distinct scenes at 11 physical sites (residential, catering, retail, office), covering 1,600 unique objects and 58 operational skills. All data uses a single embodiment — the Galaxea R1 Lite, a 23-DoF mobile bimanual robot (two 6-DoF arms, 3-DoF torso, 6-DoF omnidirectional base) — paired with precise subtask-level language annotations.

G0 dual-system model. A System 1/2 design coupling two components:

  • System-2 (G0-VLM): a Qwen2.5-VL planner (7B/32B/72B variants evaluated) fine-tuned on Galaxea subtask annotations for high-level reasoning and subtask planning.
  • System-1 (G0-VLA): initialized from PaliGemma (3B) with a SigLIP vision encoder; uses a FAST action tokenizer plus a flow-matching objective to emit continuous action chunks for fine-grained execution.

Three-stage curriculum. (1) Cross-embodiment pre-training (~1,700 h: ~1,000 h OXE + 500 h Galaxea + ~200 h in-house); (2) single-embodiment pre-training on Galaxea data; (3) task-specific post-training. The paper finds the single-embodiment stage, enabled by the Galaxea dataset, is critical to final performance. This places G0 in the VLM-planner + separate action-expert family (see VLA Architectures review) and leverages a uniform embodiment to sharpen cross-embodiment transfer.

Results

On the comprehensive benchmark (tabletop manipulation, few-shot learning, long-horizon mobile manipulation), G0 with full curriculum achieves the highest average task-progress scores, outperforming a π₀ baseline on object manipulation. In few-shot settings (20 trajectories/task), models with single-embodiment pre-training significantly outperform those without. For System-2 planning, the fine-tuned G0-VLM reaches 83.3% accuracy on Table Bussing vs. 32.0% for Gemini-2.5-pro (>50% relative gain). (Per-task low-level success rates are reported in the paper's tables; only confirmed figures are quoted here.)

Significance

Galaxea + G0 contributes a rare combination: a large single-embodiment open-world dataset with subtask language, and an explicit System-1/System-2 VLA that uses it. The key empirical message — that single-embodiment pre-training on consistent, well-annotated data drives generalization more than raw cross-embodiment scale — is a useful counterpoint for the field's cross-embodiment scaling debate and for dual-system VLA design.

Links

Related pages

← Back to ICRA-2026

⚠️ **GitHub.com Fallback** ⚠️