Review Omega0 - Heungwoo/research GitHub Wiki

In-Depth Review — ω-0: A Latent Predictive World Action Model for Concurrent Humanoid Loco-Manipulation

Paper: ω-0: A Latent Predictive World Action Model for Concurrent Humanoid Loco-Manipulation Authors: Zhe Li*†, Zhenzhe Zhang*, Yangyang Wei*, Wenjie Zhang*, Xichen Yuan*, Peiyuan Zhi, Gen Li, Xinying Guo, Fengjie Gao, Jianfei Yang♣, Shanghang Zhang♣ (* equal, † project lead, ♣ corresponding) Affiliations: MARS Lab (NTU) · Peking University · BAAI · HKUST(GZ) arXiv: 2608.06375 (v1 Aug 6, 2026, cs.RO) · Project: gentlefress.github.io/OMEGA-0_page Status: preprint (posted Aug 2026) — indexed here under Latest Papers

Companion reviews: Humanoid VLA · World Models · DreamZero · Ψ₀ · System 0/1/2 · Human Video → Robot Transfer · Real-Time Execution.


1. TL;DR

  1. A whole-body World Action Model for concurrent humanoid loco-manipulation. Most humanoid stacks decompose "move, then manipulate"; ω-0 learns a single policy that steps, leans, balances, reaches, and manipulates simultaneously — from a language instruction, current multi-view observation, and proprioceptive state — outputting controller-compatible whole-body action latents executed by the SONIC low-level controller.
  2. The key design choice: future prediction as a reconstruction-free latent objective, not video generation. Unlike video-centered humanoid WAMs (MotionWAM, DiT4DiT) that make a predicted video trajectory the intermediate representation for action, ω-0 predicts compact future observation embeddings (V-JEPA / Wan-latent targets) as a lightweight auxiliary signal, and directly denoises actions. Rationale: real humanoid vision is noisy/occluded/viewpoint-shifting during locomotion, so binding actions to a predicted video amplifies temporal inconsistencies into unstable whole-body motion — and better pixels don't imply better control.
  3. Prefix-guided dual-query attention couples foresight to action. A joint video-action latent predictor runs motion queries (one per action step) and video queries (future visual latents) with token-specific RoPE (2D for visual prefix, 3D for video queries, 1D for temporal action queries); motion queries attend to video queries so predicted scene-evolution cues are injected into the action representation — coordinated loco-manipulation without separate navigation and arm-control modules.
  4. Scales human/public motion into humanoid supervision via SONIC replay. A three-stage pipeline: (1) whole-body FAST-tokenized action VLM pretraining on Qwen3-VL-2B; (2) human-to-humanoid action-latent pretraining where public human motions (ARCTIC, Xperience-10M, Motion-X) are replayed in simulation by SONIC to produce robot-executable action latents + proprioceptive states (untrackable motions filtered out); (3) real-world fine-tuning with training-time Real-Time Chunking for smooth receding-horizon execution.
  5. Headline result: a single ω-0 model does 11 real household tasks, ~2× the best baseline. On the ω-HOME task suite, ω-0 reaches 81.8% success / 90.3% task progress vs the strongest baseline ψ-0 (44.5% / 59.6%) and video-WAM DiT4DiT (43.6% / 61.0%) — all trained on the same real data, one multi-task policy each.

2. Why this paper matters

  • It targets the axis humanoid VLAs punt on: concurrent whole-body coordination. Wiping a large table, mopping a floor, or loading a low washing-machine compartment fails if you separate locomotion, balance, and manipulation into phases. ω-0 makes whole-body coordination the learned representation, extending the Humanoid VLA frontier past the "walk there, then stand and manipulate" decomposition of Ψ-0/AMO-style stacks.
  • It is a latent WAM — the third design point vs DreamZero and MotionWAM. DreamZero jointly denoises pixel video + action in a 14B backbone; MotionWAM conditions on a video world model's denoising features. ω-0 argues that for real-time humanoid control the policy "mainly needs compact future information," so it drops pixel-space video entirely from the action path (video decode is optional visualization only). This makes it the reconstruction-free / latent-predictive corner of the WAM design space in Review-World-Models.
  • It operationalizes human-video transfer for whole-body humanoids — the human-video fork applied to legs+torso+arms, using controller-replay grounding (SONIC) to convert action-free human motion into robot-executable latents rather than co-training raw or synthesizing pixels.
  • It ships a dataset the field lacks: ω-HOME. 40.3 h / 4,827 episodes / 24 tasks at 30 Hz of household humanoid demonstrations with six synchronized modalities (egocentric RGB, exocentric RGB-D, whole-body SMPL motion, robot state, whole-body action latents, language) — closing a real gap for whole-body loco-manipulation training and evaluation.

3. Architecture

ω-0 architecture (Figure 2 of arXiv 2608.06375, © the authors)

Figure 2 of the paper. Stage 1 (left): a whole-body VLM (Qwen3-VL-2B) is fine-tuned to autoregressively predict FAST-tokenized whole-body action tokens from text + ego/exo visual tokens, giving an action-aware semantic prior. Stage 2 & 3 (center): a Joint Video-Action Latent Predictor fuses the VLM feature, T5 text, V-JEPA visual features, a view token, motion queries, and video queries; its future-aware motion feature conditions an Action DiT (0.45B) that denoises SONIC-compatible whole-body action latents, while video queries are supervised against frozen-Wan future latents. Right: the prefix-guided dual-query attention — self-attention on prefix/video/action queries, cross-attention of both query sets to the prefix, then motion queries attend to video queries — with 2D/3D/1D RoPE per token type. The predicted latents run on the SONIC whole-body controller in receding-horizon.

3.1 Problem formulation

Given language ℓ, a current view observation o^v_t, and robot state s_t, ω-0 predicts a future chunk of whole-body action latents z_{t:t+H} executed by SONIC. State is compact: joint positions q_pos, dexterous-hand joints q_hand, and torso orientation as a continuous 6D rotation (quaternion converted to 6D to avoid the double-cover discontinuity); IMU linear acceleration / angular velocity are deliberately discarded.

3.2 Three-stage training

Stage What trains Objective Data
1 · Whole-body action VLM Qwen3-VL-2B (fine-tuned) + whole-body FAST tokenizer Next-token prediction of discrete whole-body action tokens; L1 tokenizer reconstruction ARCTIC + Xperience-10M + Motion-X, unified to SMPL
2 · Human→humanoid action-latent pretraining joint predictor, state encoder, condition-fusion, Action DiT (VLM/V-JEPA/Wan frozen) x₀-prediction action denoising + L_video (future latent MSE vs frozen Wan) Public motions replayed by SONIC → robot-executable latents + states; untrackable motions filtered
3 · Real-world fine-tuning same modules Stage-2 objective + training-time RTC (random clean prefix M, loss on non-prefix only) ω-HOME real humanoid data

Why the latent objective: future visual prediction is "intentionally lightweight and only serves as an auxiliary predictive objective" — the Action DiT denoises action latents directly from language/visual/state/future-aware conditions, with no test-time video-to-action inversion and no large video generator in the loop.

3.3 Flexible viewpoints & smooth deployment

View tokens distinguish egocentric RGB, exocentric RGB, exocentric depth — exocentric RGB-D gives richer whole-body/scene supervision at training time while the robot deploys from first-person egocentric feedback. Deployment uses RTC: the last few frames of the previous chunk are cached as a clean prefix, so newly predicted chunks stay temporally consistent (the wiki's Real-Time Execution "trained-in continuation" pattern, here as a humanoid WAM).


4. ω-HOME dataset

  • 40.3 hours / 4,827 episodes / 24 tasks @ 30 Hz, household scenarios across 8 capability groups (object retrieval, surface cleaning, appliance interaction, container transfer, cloth handling, storage arrangement, mobile manipulation, tool-based floor operation).
  • Six synchronized modalities per episode: egocentric RGB (robot head cam = deployment view), exocentric RGB-D (ZED), proprioceptive state, whole-body SMPL motion references, whole-body action latents, language.
  • Teleoperation rig: Pico 4 Ultra headset + handheld triggers + foot-mounted Pico trackers → head/hand/lower-body cues retargeted to a humanoid with Inspire DexHands; SONIC serves as the teleoperation policy.
  • Leak-free protocol: the 11 downstream evaluation tasks are excluded from the Stage-2 pretraining pool; a single multi-task model is fine-tuned over all 11 tasks (no per-task policies).

5. Results

Setup: 11 real household loco-manipulation tasks (Table 1), 10 trials each, single multi-task policy per method, all trained on the same real data. Metrics: success rate, subtask score (max 41), task progress.

Method (category) Success ↑ Score/41 ↑ Task Progress ↑
ACT (classical IL) 8.2 10.6 32.4
Diffusion Policy 15.5 14.8 40.6
π0.5 (VLA) 27.3 20.9 52.8
InternVLA-M1 31.8 21.8 55.6
EgoVLA 25.5 18.6 49.1
GR00T-N1.7 22.7 19.7 49.8
ψ-0 (humanoid, arm-centric + AMO) 44.5 23.6 59.6
Fast-WAM 37.1 22.3 57.8
DiT4DiT (video WAM) 43.6 23.1 61.0
ω-0 (Ego) 79.1 35.8 88.7
ω-0 (Omni) 81.8 36.7 90.3

Reading: classical IL handles short chunks but not long-horizon whole-body coordination; VLAs gain from pretraining but their action interfaces aren't whole-body; humanoid/WAM baselines improve but are either arm-centric (ψ-0, Fast-WAM) or video-generation-centered without a controller-compatible whole-body latent interface (DiT4DiT). ω-0's ego-only variant already ~1.8× the best baseline, and adding exocentric supervision (Omni) adds a few more points. The paper reports promising held-out object/scene and human-transfer generalization qualitatively.


6. Significance & positioning

  • Sharpens the WAM taxonomy. Placed against DreamZero (pixel-video + action, 14B, arm/mobile) and MotionWAM/DiT4DiT (video-world-model-conditioned humanoid), ω-0 is the latent-predictive, reconstruction-free, whole-body corner: future prediction is a cheap auxiliary signal, not the action pathway. It's a concrete counter-argument to "improving robotics = improving video generation" (DreamZero's thesis) for the humanoid real-time regime, where the authors argue compact future info beats pixel fidelity.
  • Controller-grounded human data is the scalable lever. SONIC replay converts abundant action-free human/public motion into robot-executable whole-body latents — a different resolution of the human-video fork than emergence/decoupling/pixel-synthesis, specific to whole-body humanoids.
  • A missing dataset filled. ω-HOME's whole-body, multi-view, SMPL-annotated household corpus is directly reusable and addresses the "humanoid evaluation lags arms by a generation" gap noted in Review-Humanoid-VLA.

7. Limitations

7.1 Visible in the paper

  • Depends on the SONIC controller. ω-0 predicts SONIC-compatible latents and uses SONIC for both teleoperation and low-level execution; portability to other whole-body controllers is untested, and the action interface is defined by SONIC's trackable-motion distribution (untrackable motions are filtered out of training).
  • Single embodiment / lab-scale human transfer. Results are on one humanoid; human-transfer generalization is shown but at small scale, and cross-embodiment robustness is qualitative.
  • Absolute long-horizon success is still moderate. 81.8% is a large relative win, but on a bespoke 11-task suite with a self-defined subtask/progress protocol; no shared external humanoid benchmark exists to place it against other labs.

7.2 Reviewer's notes

  • The reconstruction-free thesis is argued and supported on this setup, but there's no head-to-head ablation isolating "latent future target vs pixel-video target" at matched scale on the same tasks — the case against video-centered WAMs is made partly by baseline comparison rather than a controlled swap.
  • IMU dynamics (accel/angular vel) are discarded for stability; whether that caps highly dynamic behaviors is unexplored.
  • As a very recent preprint (Aug 2026), numbers are v1 and unreplicated; treat as early-signal.

8. Links & related pages

← Back to Latest Papers · Home · Reviews