CVPR 2026 GigaBrain 0.5M - Heungwoo/research GitHub Wiki

GigaBrain-0.5M* โ€” a VLA That Learns From World-Model-Based RL

Venue: CVPR 2026 Category: World-Model / RL-augmented VLA Trend tag: Trend 3 Team: GigaAI (GigaBrain Team)

Approach diagram

flowchart TB
  WM["pretrained video world model<br/>(future + value prediction)"] --> COND["condition policy on<br/>predicted futures + value estimates"]
  COND --> FT["fine-tune VLA policy"]
  FT --> DEPLOY["deploy with human intervention<br/>collect trajectories"]
  DEPLOY --> REFINE["jointly refine policy + world model<br/>on curated rollouts"]
  REFINE --> WM
Loading

Problem

RL-augmented VLAs (VLA-RFT, RECAP) treat reward as a scalar signal. But a learned world model contains far richer information: value over future states, future-state predictions, dynamics consistency. Standard RL pipelines do not exploit this.

Method

RAMP (Reinforcement leArning via world Model-conditioned Policy) โ€” a four-stage iterative training paradigm built on the GigaBrain-0.5 base model (pretrained on 10,000+ hours of robotic manipulation data; ~61% synthesized by GigaWorld, ~39% real-robot). Rather than reducing the world model to a scalar reward (as in sparse-advantage RL) or to a DreamGen-style imagination buffer, RAMP conditions the policy on the world model's rich predicted futures and value estimates during policy fine-tuning. The loop: (1) pretrain a video world model on manipulation data; (2) fine-tune the policy with actions conditioned on predicted futures + value estimates; (3) deploy in the physical environment with human intervention to collect trajectories; (4) jointly refine policy and world model on curated rollout data, then iterate. The conditioning is part of the RL training loop, not an inference-time planner.

Results

~30 % improvement over the RECAP baseline on challenging long-horizon tasks โ€” Laundry Folding, Box Packing, Espresso Preparation โ€” with reliable long-horizon execution. The base GigaBrain-0.5 intermediate model also ranks first on the international RoboChallenge benchmark.

Significance

GigaBrain-0.5M*'s contribution is using the world model's dense conditioning signal (predicted futures + value estimates) to drive RL, rather than collapsing it to a sparse scalar advantage. It descends from GigaBrain-0 (world-model-generated data for VLA pretraining) and the GigaWorld data engine, and sits adjacent to Ctrl-World (controllable world model as RL substrate) and DreamGen (world model as training environment). RAMP is the model-conditioned-RL complement.

Links

Related pages

โ† Back to CVPR-2026

โš ๏ธ **GitHub.com Fallback** โš ๏ธ