CVPR 2026 GigaBrain 0.5M - Heungwoo/research GitHub Wiki
Venue: CVPR 2026 Category: World-Model / RL-augmented VLA Trend tag: Trend 3 Team: GigaAI (GigaBrain Team)
flowchart TB
WM["pretrained video world model<br/>(future + value prediction)"] --> COND["condition policy on<br/>predicted futures + value estimates"]
COND --> FT["fine-tune VLA policy"]
FT --> DEPLOY["deploy with human intervention<br/>collect trajectories"]
DEPLOY --> REFINE["jointly refine policy + world model<br/>on curated rollouts"]
REFINE --> WM
RL-augmented VLAs (VLA-RFT, RECAP) treat reward as a scalar signal. But a learned world model contains far richer information: value over future states, future-state predictions, dynamics consistency. Standard RL pipelines do not exploit this.
RAMP (Reinforcement leArning via world Model-conditioned Policy) โ a four-stage iterative training paradigm built on the GigaBrain-0.5 base model (pretrained on 10,000+ hours of robotic manipulation data; ~61% synthesized by GigaWorld, ~39% real-robot). Rather than reducing the world model to a scalar reward (as in sparse-advantage RL) or to a DreamGen-style imagination buffer, RAMP conditions the policy on the world model's rich predicted futures and value estimates during policy fine-tuning. The loop: (1) pretrain a video world model on manipulation data; (2) fine-tune the policy with actions conditioned on predicted futures + value estimates; (3) deploy in the physical environment with human intervention to collect trajectories; (4) jointly refine policy and world model on curated rollout data, then iterate. The conditioning is part of the RL training loop, not an inference-time planner.
~30 % improvement over the RECAP baseline on challenging long-horizon tasks โ Laundry Folding, Box Packing, Espresso Preparation โ with reliable long-horizon execution. The base GigaBrain-0.5 intermediate model also ranks first on the international RoboChallenge benchmark.
GigaBrain-0.5M*'s contribution is using the world model's dense conditioning signal (predicted futures + value estimates) to drive RL, rather than collapsing it to a sparse scalar advantage. It descends from GigaBrain-0 (world-model-generated data for VLA pretraining) and the GigaWorld data engine, and sits adjacent to Ctrl-World (controllable world model as RL substrate) and DreamGen (world model as training environment). RAMP is the model-conditioned-RL complement.
- arXiv: 2602.12099
โ Back to CVPR-2026