Review Omega0 - Heungwoo/research GitHub Wiki
In-Depth Review — ω-0: A Latent Predictive World Action Model for Concurrent Humanoid Loco-Manipulation
Paper: ω-0: A Latent Predictive World Action Model for Concurrent Humanoid Loco-Manipulation Authors: Zhe Li*†, Zhenzhe Zhang*, Yangyang Wei*, Wenjie Zhang*, Xichen Yuan*, Peiyuan Zhi, Gen Li, Xinying Guo, Fengjie Gao, Jianfei Yang♣, Shanghang Zhang♣ (* equal, † project lead, ♣ corresponding) Affiliations: MARS Lab (NTU) · Peking University · BAAI · HKUST(GZ) arXiv: 2608.06375 (v1 Aug 6, 2026, cs.RO) · Project: gentlefress.github.io/OMEGA-0_page Status: preprint (posted Aug 2026) — indexed here under Latest Papers
Companion reviews: Humanoid VLA · World Models · DreamZero · Ψ₀ · System 0/1/2 · Human Video → Robot Transfer · Real-Time Execution.
1. TL;DR
- A whole-body World Action Model for concurrent humanoid loco-manipulation. Most humanoid stacks decompose "move, then manipulate"; ω-0 learns a single policy that steps, leans, balances, reaches, and manipulates simultaneously — from a language instruction, current multi-view observation, and proprioceptive state — outputting controller-compatible whole-body action latents executed by the SONIC low-level controller.
- The key design choice: future prediction as a reconstruction-free latent objective, not video generation. Unlike video-centered humanoid WAMs (MotionWAM, DiT4DiT) that make a predicted video trajectory the intermediate representation for action, ω-0 predicts compact future observation embeddings (V-JEPA / Wan-latent targets) as a lightweight auxiliary signal, and directly denoises actions. Rationale: real humanoid vision is noisy/occluded/viewpoint-shifting during locomotion, so binding actions to a predicted video amplifies temporal inconsistencies into unstable whole-body motion — and better pixels don't imply better control.
- Prefix-guided dual-query attention couples foresight to action. A joint video-action latent predictor runs motion queries (one per action step) and video queries (future visual latents) with token-specific RoPE (2D for visual prefix, 3D for video queries, 1D for temporal action queries); motion queries attend to video queries so predicted scene-evolution cues are injected into the action representation — coordinated loco-manipulation without separate navigation and arm-control modules.
- Scales human/public motion into humanoid supervision via SONIC replay. A three-stage pipeline: (1) whole-body FAST-tokenized action VLM pretraining on Qwen3-VL-2B; (2) human-to-humanoid action-latent pretraining where public human motions (ARCTIC, Xperience-10M, Motion-X) are replayed in simulation by SONIC to produce robot-executable action latents + proprioceptive states (untrackable motions filtered out); (3) real-world fine-tuning with training-time Real-Time Chunking for smooth receding-horizon execution.
- Headline result: a single ω-0 model does 11 real household tasks, ~2× the best baseline. On the ω-HOME task suite, ω-0 reaches 81.8% success / 90.3% task progress vs the strongest baseline ψ-0 (44.5% / 59.6%) and video-WAM DiT4DiT (43.6% / 61.0%) — all trained on the same real data, one multi-task policy each.
2. Why this paper matters
- It targets the axis humanoid VLAs punt on: concurrent whole-body coordination. Wiping a large table, mopping a floor, or loading a low washing-machine compartment fails if you separate locomotion, balance, and manipulation into phases. ω-0 makes whole-body coordination the learned representation, extending the Humanoid VLA frontier past the "walk there, then stand and manipulate" decomposition of Ψ-0/AMO-style stacks.
- It is a latent WAM — the third design point vs DreamZero and MotionWAM. DreamZero jointly denoises pixel video + action in a 14B backbone; MotionWAM conditions on a video world model's denoising features. ω-0 argues that for real-time humanoid control the policy "mainly needs compact future information," so it drops pixel-space video entirely from the action path (video decode is optional visualization only). This makes it the reconstruction-free / latent-predictive corner of the WAM design space in Review-World-Models.
- It operationalizes human-video transfer for whole-body humanoids — the human-video fork applied to legs+torso+arms, using controller-replay grounding (SONIC) to convert action-free human motion into robot-executable latents rather than co-training raw or synthesizing pixels.
- It ships a dataset the field lacks: ω-HOME. 40.3 h / 4,827 episodes / 24 tasks at 30 Hz of household humanoid demonstrations with six synchronized modalities (egocentric RGB, exocentric RGB-D, whole-body SMPL motion, robot state, whole-body action latents, language) — closing a real gap for whole-body loco-manipulation training and evaluation.
3. Architecture

Figure 2 of the paper. Stage 1 (left): a whole-body VLM (Qwen3-VL-2B) is fine-tuned to autoregressively predict FAST-tokenized whole-body action tokens from text + ego/exo visual tokens, giving an action-aware semantic prior. Stage 2 & 3 (center): a Joint Video-Action Latent Predictor fuses the VLM feature, T5 text, V-JEPA visual features, a view token, motion queries, and video queries; its future-aware motion feature conditions an Action DiT (0.45B) that denoises SONIC-compatible whole-body action latents, while video queries are supervised against frozen-Wan future latents. Right: the prefix-guided dual-query attention — self-attention on prefix/video/action queries, cross-attention of both query sets to the prefix, then motion queries attend to video queries — with 2D/3D/1D RoPE per token type. The predicted latents run on the SONIC whole-body controller in receding-horizon.
3.1 Problem formulation
Given language ℓ, a current view observation o^v_t, and robot state s_t, ω-0 predicts a future chunk of whole-body action latents z_{t:t+H} executed by SONIC. State is compact: joint positions q_pos, dexterous-hand joints q_hand, and torso orientation as a continuous 6D rotation (quaternion converted to 6D to avoid the double-cover discontinuity); IMU linear acceleration / angular velocity are deliberately discarded.
3.2 Three-stage training
| Stage | What trains | Objective | Data |
|---|---|---|---|
| 1 · Whole-body action VLM | Qwen3-VL-2B (fine-tuned) + whole-body FAST tokenizer | Next-token prediction of discrete whole-body action tokens; L1 tokenizer reconstruction | ARCTIC + Xperience-10M + Motion-X, unified to SMPL |
| 2 · Human→humanoid action-latent pretraining | joint predictor, state encoder, condition-fusion, Action DiT (VLM/V-JEPA/Wan frozen) | x₀-prediction action denoising + L_video (future latent MSE vs frozen Wan) |
Public motions replayed by SONIC → robot-executable latents + states; untrackable motions filtered |
| 3 · Real-world fine-tuning | same modules | Stage-2 objective + training-time RTC (random clean prefix M, loss on non-prefix only) | ω-HOME real humanoid data |
Why the latent objective: future visual prediction is "intentionally lightweight and only serves as an auxiliary predictive objective" — the Action DiT denoises action latents directly from language/visual/state/future-aware conditions, with no test-time video-to-action inversion and no large video generator in the loop.
3.3 Flexible viewpoints & smooth deployment
View tokens distinguish egocentric RGB, exocentric RGB, exocentric depth — exocentric RGB-D gives richer whole-body/scene supervision at training time while the robot deploys from first-person egocentric feedback. Deployment uses RTC: the last few frames of the previous chunk are cached as a clean prefix, so newly predicted chunks stay temporally consistent (the wiki's Real-Time Execution "trained-in continuation" pattern, here as a humanoid WAM).
4. ω-HOME dataset
- 40.3 hours / 4,827 episodes / 24 tasks @ 30 Hz, household scenarios across 8 capability groups (object retrieval, surface cleaning, appliance interaction, container transfer, cloth handling, storage arrangement, mobile manipulation, tool-based floor operation).
- Six synchronized modalities per episode: egocentric RGB (robot head cam = deployment view), exocentric RGB-D (ZED), proprioceptive state, whole-body SMPL motion references, whole-body action latents, language.
- Teleoperation rig: Pico 4 Ultra headset + handheld triggers + foot-mounted Pico trackers → head/hand/lower-body cues retargeted to a humanoid with Inspire DexHands; SONIC serves as the teleoperation policy.
- Leak-free protocol: the 11 downstream evaluation tasks are excluded from the Stage-2 pretraining pool; a single multi-task model is fine-tuned over all 11 tasks (no per-task policies).
5. Results
Setup: 11 real household loco-manipulation tasks (Table 1), 10 trials each, single multi-task policy per method, all trained on the same real data. Metrics: success rate, subtask score (max 41), task progress.
| Method (category) | Success ↑ | Score/41 ↑ | Task Progress ↑ |
|---|---|---|---|
| ACT (classical IL) | 8.2 | 10.6 | 32.4 |
| Diffusion Policy | 15.5 | 14.8 | 40.6 |
| π0.5 (VLA) | 27.3 | 20.9 | 52.8 |
| InternVLA-M1 | 31.8 | 21.8 | 55.6 |
| EgoVLA | 25.5 | 18.6 | 49.1 |
| GR00T-N1.7 | 22.7 | 19.7 | 49.8 |
| ψ-0 (humanoid, arm-centric + AMO) | 44.5 | 23.6 | 59.6 |
| Fast-WAM | 37.1 | 22.3 | 57.8 |
| DiT4DiT (video WAM) | 43.6 | 23.1 | 61.0 |
| ω-0 (Ego) | 79.1 | 35.8 | 88.7 |
| ω-0 (Omni) | 81.8 | 36.7 | 90.3 |
Reading: classical IL handles short chunks but not long-horizon whole-body coordination; VLAs gain from pretraining but their action interfaces aren't whole-body; humanoid/WAM baselines improve but are either arm-centric (ψ-0, Fast-WAM) or video-generation-centered without a controller-compatible whole-body latent interface (DiT4DiT). ω-0's ego-only variant already ~1.8× the best baseline, and adding exocentric supervision (Omni) adds a few more points. The paper reports promising held-out object/scene and human-transfer generalization qualitatively.
6. Significance & positioning
- Sharpens the WAM taxonomy. Placed against DreamZero (pixel-video + action, 14B, arm/mobile) and MotionWAM/DiT4DiT (video-world-model-conditioned humanoid), ω-0 is the latent-predictive, reconstruction-free, whole-body corner: future prediction is a cheap auxiliary signal, not the action pathway. It's a concrete counter-argument to "improving robotics = improving video generation" (DreamZero's thesis) for the humanoid real-time regime, where the authors argue compact future info beats pixel fidelity.
- Controller-grounded human data is the scalable lever. SONIC replay converts abundant action-free human/public motion into robot-executable whole-body latents — a different resolution of the human-video fork than emergence/decoupling/pixel-synthesis, specific to whole-body humanoids.
- A missing dataset filled. ω-HOME's whole-body, multi-view, SMPL-annotated household corpus is directly reusable and addresses the "humanoid evaluation lags arms by a generation" gap noted in Review-Humanoid-VLA.
7. Limitations
7.1 Visible in the paper
- Depends on the SONIC controller. ω-0 predicts SONIC-compatible latents and uses SONIC for both teleoperation and low-level execution; portability to other whole-body controllers is untested, and the action interface is defined by SONIC's trackable-motion distribution (untrackable motions are filtered out of training).
- Single embodiment / lab-scale human transfer. Results are on one humanoid; human-transfer generalization is shown but at small scale, and cross-embodiment robustness is qualitative.
- Absolute long-horizon success is still moderate. 81.8% is a large relative win, but on a bespoke 11-task suite with a self-defined subtask/progress protocol; no shared external humanoid benchmark exists to place it against other labs.
7.2 Reviewer's notes
- The reconstruction-free thesis is argued and supported on this setup, but there's no head-to-head ablation isolating "latent future target vs pixel-video target" at matched scale on the same tasks — the case against video-centered WAMs is made partly by baseline comparison rather than a controlled swap.
- IMU dynamics (accel/angular vel) are discarded for stability; whether that caps highly dynamic behaviors is unexplored.
- As a very recent preprint (Aug 2026), numbers are v1 and unreplicated; treat as early-signal.
8. Links & related pages
- arXiv: https://arxiv.org/abs/2608.06375 · Project: https://gentlefress.github.io/OMEGA-0_page/
- Humanoid VLA — the frontier this extends to concurrent whole-body coordination
- World Models — the WAM taxonomy; ω-0 is the latent-predictive corner
- DreamZero — the pixel-video WAM counterpoint (and its "improve video → improve policy" thesis)
- Ψ₀ — the staged-training humanoid baseline (ψ-0) it more-than-doubles
- Human Video → Robot Transfer — controller-replay grounding as a fourth transfer route
- Real-Time Execution — training-time RTC for smooth receding-horizon control
- Latest Papers — the preprint tracker this is filed under
← Back to Latest Papers · Home · Reviews