IROS 2025 AgiBot World - Heungwoo/research GitHub Wiki

AgiBot World Colosseo — Million-Trajectory Manipulation Platform

Venue: IROS 2025 (Best Paper Finalist) · Authors: AgiBot-World Contributors (AgiBot / Shanghai AI Lab et al.) · arXiv: 2503.06669 Category: Data / Benchmark (large-scale real-robot dataset) Trend tag: Scaling real-world manipulation data

The largest open real-world manipulation corpus to date: over 1 million trajectories across 217 tasks in five deployment domains, collected on ~100 real bimanual robots. The platform ships not just data but a foundation policy (GO-1, Genie Operator-1) and the ViLLA training framework. The repo lists it as IROS 2025 Best Paper Award Finalist & IEEE T-RO 2026.

Approach diagram

flowchart LR
  Robots["~100 real robots<br/>(AgiBot G1, bimanual)"] --> Domains["5 domains<br/>domestic · retail · industrial<br/>restaurant · office"]
  Domains --> Sensing["Multimodal sensing<br/>visuo-tactile · 6-DoF dexterous hands<br/>long-horizon / durative actions"]
  Sensing --> Pipeline["Standardized collection<br/>+ human-in-the-loop QC"]
  Pipeline --> Data["1M+ trajectories<br/>217 tasks"]
  Data --> ViLLA["ViLLA framework<br/>LAM → Latent Planner → Action Expert"]
  ViLLA --> GO1["GO-1 generalist policy<br/>latent-action foundation model"]
Loading

Problem

Generalized robotic manipulation is bottlenecked by data: existing public datasets (Open X-Embodiment, DROID, BridgeData V2) are heterogeneous and small relative to what scaling laws in other modalities suggest is needed. The paper asks whether scalable, high-quality real-world robot data can be collected at order-of-magnitude larger scale, and whether policies trained on it show predictable performance scaling with data volume.

Method — Platform & dataset

  • Dataset. Over 1 million trajectories spanning 217 tasks in five deployment scenarios: domestic, retail, industrial, restaurant, and office. An order-of-magnitude increase in scale over prior public datasets. Trajectories typically span ~30 seconds, some over 2 minutes (long-horizon / durative actions).
  • Hardware. Collected on the AgiBot G1 platform — ~100 homogeneous real bimanual robots. Extensible from grippers to 6-DoF dexterous hands and visuo-tactile sensors for fine-grained skill acquisition.
  • Collection pipeline. Standardized teleoperation collection with human-in-the-loop verification to guarantee high-quality, diverse data distribution.
  • GO-1 (Genie Operator-1). A generalist policy built on the ViLLA (hierarchical Vision-Language-Latent-Action) framework, which decouples planning from low-level control via latent actions:
    1. Latent Action Model (LAM) — encoder-decoder projecting consecutive images into a latent action space, trainable on internet-scale heterogeneous video.
    2. Latent Planner — a pretrained VLM doing embodiment-agnostic planning in the latent action space.
    3. Action Expert — a diffusion objective modeling the continuous distribution of low-level actions for high-frequency, dexterous control.

Results

  • Policies pre-trained on AgiBot World achieve an average +30% performance over those trained on Open X-Embodiment, both in-domain and out-of-distribution.
  • GO-1 reaches over 60% success on complex real-world dexterous and long-horizon tasks, outperforming the prior RDT approach by 32%.
  • GO-1 demonstrates predictable performance scaling with increased data volume — the core scaling-law claim of the platform.

(Figures above are the only quantitative claims stated in the abstract; per-task breakdowns are in the paper.)

Significance

AgiBot World is the real-world data-scaling counterpart to the simulation- and benchmark-scaling threads tracked elsewhere in this wiki. Where ML-venue work optimizes architecture and training recipes, AgiBot World argues the binding constraint is data, and provides the substrate: it is cited across the IROS/data thread as "the largest public manipulation dataset." It sits alongside the deployment-and-hardware framing of the IROS 2025 Survey, and its latent-action / cross-embodiment design connects to the generalization arguments in the Cross-Embodiment review. Distinguish the dataset/platform (the corpus + collection pipeline + AgiBot G1 hardware) from the GO-1 model (the ViLLA-based foundation policy trained on it) — they are released together but are separable contributions.

Links

Related pages

← Back to Home

⚠️ **GitHub.com Fallback** ⚠️