Review Being H07 - Heungwoo/research GitHub Wiki
Model: Being-H0.7 — latent world-action model (unified and reactive WAM+VLA hybrid) · BeingBeyond (Zongqing Lu group) Paper: arXiv 2605.00078 (Apr 30 2026) · project page · GitHub BeingBeyond/Being-H Status: arXiv preprint (not yet peer-reviewed) — filed under Latest Papers. Authors: Hao Luo, Wanpeng Zhang, Yicheng Feng, Sipeng Zheng, … Zongqing Lu. Built on the pretrained VLA Being-H0.5.
The blog's "most sophisticated hybrid" (NVIDIA WAM thesis §2.6): it fuses world-model imagination with VLA deployability by predicting a latent reasoning state rather than pixels — landing in both hybrid corners at once (unified architecture, reactive inference). Companion: VLA Architectures §4.2b · World Models · DYNA-2 · ω-0 · MOTUS.
- Latent queries as an explicit reasoning interface. Being-H0.7 inserts a small set of learnable latent queries between the multimodal context and the action tokens. These slots attend to instruction + observation history + robot state and form a compact latent state before actions are generated — a "think, then act" bottleneck without generating any future frames.
- Dual-branch posterior/prior — the key trick. A future-informed posterior branch (training-only) replaces the queries with embeddings from future observations; a deployable prior branch infers the latent state from the current context alone. Aligning the two teaches the prior to encode future-aware, action-useful structure from the present — with hidden-state alignment + regularization to prevent latent collapse.
- Reactive at inference. At deployment it discards the posterior branch and performs no visual rollout — like DYNA-2/ω-0, it keeps the world-model benefit while staying a fast VLA-style policy.
- Backbones: understanding = InternVL3.5, action = Qwen3, visual encoders = V-JEPA 2.1 (context-frame encoder kept trainable). Pretrained on large-scale egocentric human video + robot demonstrations (reported ~200k h human + ~15k h robot).
- It reconciles imagination and deployability without pixels. Rather than choose between an in-path video WAM (strong prior, slow) and a bare VLA (fast, weak prior), Being-H0.7 predicts a latent future-informed state and drops the predictive machinery at inference — the cleanest "reactive latent WAM" recipe alongside ω-0.
-
Posterior/prior is a principled version of "co-training dropped at inference." Where DYNA-2 simply omits
z_tfrom the action head, Being-H0.7 explicitly trains a prior to match a future-informed posterior — a Play-LMP-style latent-plan structure that gives the reactive bet a training objective. - Egocentric-video pretraining at scale, continuing the human-video fork that DYNA-2 and EgoScale push.
flowchart LR
CTX[instruction · obs history · robot state] --> Q[Learnable latent queries<br/>compact reasoning state]
Q --> A[Action tokens · Qwen3]
subgraph TRAIN[training only]
FUT[future observations] --> POST[Posterior branch<br/>queries ← future embeddings]
end
POST -. align latent reasoning space<br/>+ anti-collapse reg .- Q
Q -->|deployable prior: no rollout| A
- Understanding tower: InternVL3.5. Action tower: Qwen3. Visual encoders: V-JEPA 2.1 (both), with the context-frame encoder trainable.
- Prior (deployable): infers latent reasoning state from current context only.
- Posterior (training-only): substitutes the latent queries with embeddings from future observations; the two branches are jointly aligned in latent reasoning space, with hidden-state alignment + lightweight regularization to avoid latent collapse.
- Inference: posterior discarded, no visual rollout — reactive.
| Benchmark | Being-H0.7 |
|---|---|
| LIBERO | 99.2% |
| LIBERO-plus (zero-shot / fine-tuned) | 82.1% / 84.8% |
| RoboTwin 2.0 Hard | 89.6% |
| RoboCasa-50 | 62.1% |
| GR1 | 49.2% |
| CALVIN (ABCD→D / ABC→D) | 4.67 / 4.48 tasks |
- Outperforms π0.5, Fast-WAM, and Being-H0.5 across most benchmarks.
- Real world: leads on all five ability-oriented task suites across three platforms — PND Adam-U, Unitree G1, Franka FR3.
Significance. Being-H0.7 is the strongest academic case that a latent, reactive world-action model can match in-path WAMs on task quality while deploying like a VLA — and it gives the "reactive" bet (DYNA-2, ω-0) a principled posterior/prior training objective rather than a design omission. It sits at the intersection of both hybrid axes in Review-VLA-Architecture §4.2b (unified architecture, reactive inference).
Limitations.
- arXiv preprint, unreplicated; numbers are the authors' own.
- Action-generation focus. The paper notes it currently focuses on action generation rather than text-generation tasks — the VLM's language breadth under this bottleneck isn't stressed.
- Latent-collapse risk is designed around, not eliminated — the regularization is a mitigation; robustness of the prior/posterior gap across domains isn't fully characterized.
- Exact pretraining hours are reported at the abstract level (~200k h human + ~15k h robot); the per-source breakdown is less detailed than the benchmark tables.
- Paper: arXiv 2605.00078 · project: research.beingbeyond.com/being-h07 · code: GitHub
- VLA Architectures §4.2b (WAM×VLA hybrids) · World Models · Human Video → Robot Transfer
- Sibling hybrids: DYNA-2 · ω-0 · MOTUS · Cortex 2.0 · Cosmos 3 / NVIDIA WAM
- Latest Papers
← Back to Latest Papers · Home · Reviews