ICLR 2026 Genie Envisioner - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 Category: World Model / VLA Foundation Trend tag: Trend 8 (world models) + Trend 1 (foundation-scale data) Authors: AgiBot Genie Team + LV-NUS Lab + BUAA (Beihang). Yue Liao, Pengfei Zhou, Siyuan Huang, Donglin Yang, Shengcong Chen, et al.; Shuicheng Yan and Maoqing Yao corresponding.
flowchart LR
subgraph GEBase[GE-Base · LTX-Video 2B DiT · multi-view video diffusion]
Inst[Instruction q + initial multi-view obs x0] --> Enc[Shared video encoder + RoPE + view embedding]
Mem[Sparse memory frames m_0:t-1] --> Enc
Enc --> XV[Cross-view attention<br/>in α·N blocks]
XV --> Lat[Multi-scale latent features v_i<br/>across DiT depth]
end
Lat --> GEAct[GE-Act · 160M parallel action DiT<br/>block-wise aligned with GE-Base]
GEAct --> CA[Cross-attend to v_i at every block]
CA --> FM[Flow-matching velocity v^act_θ]
FM --> Act[54-step action chunk @ 30 Hz]
Act --> Robot[AgiBot G1 / Dual Franka / Agilex Cobot Magic / RoboTwin]
GEBase -.5 Hz video refresh.- Async[Slow-Fast asynchronous inference]
GEAct -.30 Hz action.- Async
Most manipulation stacks split visual prediction (video diffusion) and action policy learning into separate pipelines, leading to awkward interfaces, high inference latency, and loss of spatial detail. The paper identifies two specific failure modes in the prior literature:
- Language-centric VLAs (RT-2, OpenVLA, π0) compress visual observations into low-bandwidth semantic embeddings that fail to encode future dynamics, limiting fine-grained motor control. Mixing diffusion-based action losses with language objectives further destabilizes pretrained weights.
- Video-as-policy backbones (VPP, UVA, Vid-MAN) typically use a serial video-to-action pipeline that compresses video latents before policy decoding, discarding fine-grained motion and contact cues. Pixel-level action decoding compounds this with long inference latency.
Genie Envisioner argues for a single foundation that unifies world modeling and policy in one closed-loop generative architecture.
Backbone. GE-Base extends the LTX-Video 2B DiT (HaCohen et al., 2024) into an autoregressive chunk-wise multi-view video generator. At step t, it predicts a chunk of N consecutive frames xt1:N conditioned on:
- initial multi-view observation x0 (head + left wrist + right wrist),
- sparse memory frames m0:t-1 from previously generated chunks (long-horizon context),
- instruction embedding T(q).
Per-view tokens are encoded via a shared VAE encoder E, augmented with a 3D rotary positional embedding (time, height, width) and a learnable view embedding eiview:
ṽi = RoPE(t, h, w) + vi + eiview
Per-view inputs ui = ṽ0i ∥ ṽmi ∥ zi (where zi is view-specific Gaussian noise) are concatenated across views and fed to the DiT.
Cross-view attention. A subset α·Nblocks of the DiT blocks fuse all views into one latent sequence (cross-view self-attention) for geometric coherence; the remaining (1−α)·Nblocks process views independently to save compute (Figure 3b). This is the "causal block" design.
Latent diffusion training objective. With VAE latent l of the target chunk, noised latent l̃ = (1−στ)l + στε, the model predicts a denoising velocity vθ:
Lvideo = w(τ) ‖ vθ − (ε − l) ⊙ (1 − M) ‖²
where M is a conditioning mask so supervision falls only on future frames.
- Data: AgiBot-World-Beta — ~1M dual-arm manipulation episodes with three calibrated cameras and language annotations (Bu et al., 2025a). The abstract rounds this to "~3,000 hours"; §2.2 gives the precise figure of 2,967 hours of paired video–language demonstrations.
- Stage I: Multi-Resolution Temporal Adaptation (GE-Base-MR). 57-frame clips sampled at 3–30 Hz, plus 4 sparse memory frames; clips compressed into an 8-frame latent space via a frozen video VAE. Trains spatiotemporal representations invariant to sampling rate. Compute: 7 days × 32 GPUs.
- Stage II: Low-Frequency Policy Alignment (GE-Base-LF). Fine-tune on 9-frame clips at 5 Hz with 4 memory frames; encoded to 2 latent frames; only generation components are updated. Aligns temporal abstraction with downstream control. Compute: 3 days × 32 GPUs.
Architecture. GE-Act mirrors GE-Base's DiT depth but with reduced hidden dimension — a 160M-parameter lightweight autoregressive decoder. At each DiT block depth i it cross-attends to the corresponding multi-scale visual feature vi:
ai = Biact(zact, CrossAttn(zact, vi))
This block-wise alignment is the paper's core architectural bet: instead of conditioning policy on the final video latent (as serial VPP/UVA do), GE-Act consumes high-resolution latents from every DiT depth, preserving fine-grained spatial cues and cross-view correspondences.
Memory difference vs GE-Base. GE-Act samples memory directly from real robot observations, not from generated frames — so action conditioning never drifts on hallucinated history.
Training (3 stages, Figure 4 right).
- Action pre-training on AgiBot-World-Beta: GE-Base-LF frozen; only the action decoder updates. Video generation disabled; 4 memory frames at 5 Hz; predict 54-step action chunks at 30 Hz. 3 days × 16 GPUs.
- Stage 2: Task-specific video adaptation — only video components updated on a composite of AgiBot-World + task-specific data. 12 hours × 8 GPUs.
- Stage 3: Task-specific action specialization — full model fine-tuned. ~36 hours × 8 GPUs.
Loss. Velocity-matching identical in form to GE-Base:
Lact = w(τ) ‖ vθact − (ε − u) ‖²
Two independent forms of asynchrony:
- Diffusion-step asynchrony: the heavy video DiT does single-step denoising when refreshing latents; the action decoder does multi-step denoising.
- Frequency asynchrony: GE-Act Slow updates both branches at the same rate; GE-Act Fast updates video at 5 Hz and actions at 30 Hz.
GE-Act Fast supports a 54-step prediction window and 30 action steps within 200 ms on an RTX 4090.
GE-Base vs general video models in the unified text-and-image-to-video setting, evaluated on the EWMBench Scene / Motion / Semantics dimensions. The paper presents this as Figure 17, which combines a fine-grained radar plot (a) with an aggregated numeric table (b):
| Model | Scene | Motion | Semantics | Score |
|---|---|---|---|---|
| GE-Base | 0.9427 | 1.6676 | 2.0907 | 4.7010 |
| Kling | 0.8888 | 0.9440 | 2.0370 | 3.8698 |
| Hailuo | 0.8577 | 0.5362 | 2.0186 | 3.4125 |
| COSMOS | 0.7963 | 0.7085 | 1.7824 | 3.2872 |
| OpenSora | 0.9210 | 0.3442 | 1.8739 | 3.1392 |
| LTX | 0.9156 | 0.4002 | 1.6518 | 2.9676 |
(Source: Figure 17(b) "Aggregated Evaluation Across Hierarchical Levels", arXiv:2508.05635 — values verified against the PDF.) GE-Base dominates on Scene and Motion; Semantics is comparable to general video models — the paper attributes this to embodied-data pretraining capturing task-relevant spatiotemporal dynamics that generic video models lack. The text (§6.3) states GE-Base "consistently outperforms the baselines across multiple evaluation dimensions," with notable strengths in temporal alignment and dynamic consistency. The accompanying human-preference study (Figure 18) shows EWMBench rankings concord with human judgments.
Audit note (round 3, 2026-06-01): a round-1 edit had removed this table, asserting the numbers were Table 2 (EnerVerse_FT) of the separate EWMBench benchmark paper (arXiv:2505.09694) and that no "GE-Base" row existed. That was a WebFetch truncation false-positive: the full PDF of the Genie Envisioner paper (arXiv:2508.05635) does contain this exact table as Figure 17(b), with the GE-Base row at 0.9427 / 1.6676 / 2.0907 / 4.7010 and the Kling/Hailuo/COSMOS/OpenSora/LTX rows reproduced above verbatim. The table has been restored. (The original wiki heading had mislabeled it "Figure 5"; the correct figure number is 17.)
5 tasks (sandwich, pour tea, clean table, microwave, conveyor packing). Compared against π0, UniVLA, GR00T N1 under identical fine-tuning data. Two metrics: Step-wise Success Rate (SR) and End-to-End Success Rate (E2E). GE-Act-Slow and GE-Act-Fast both consistently surpass baselines; asynchronous mode is comparable or better, particularly on latency-sensitive tasks (dynamic tracking) and short-horizon tasks (packing detergent). Per-task numerical scores are presented as a bar chart in Figure 7 (not as a table; precise per-task numbers not stated in the body text).
Dual Franka cloth folding (Figure 8a): 250 teleoperated episodes (1 hour) on a space-mouse interface. GE-Act outperforms π0, UniVLA, GR00T N1 — even though π0/GR00T N1 saw large-scale Franka data during pretraining and GE-Act did not.
Agilex Cobot Magic (Figure 8b): 250 demos (~1 hour) per task using Aloha-style teleop. Two tasks: box folding and cloth folding. GE-Act consistently outperforms π0, UniVLA, GO-1, GR00T N1. UniVLA and GR00T N1 achieve 0% success on complex deformable tasks; π0 shows stronger deformable performance but GE-Act significantly surpasses it.
RoboTwin simulator (Figure 8c): All-in-one fine-tuning on 4 tasks with 200 demos (50/task). Single unified GE-Act model evaluated on all four tasks (grab roller, handover mic, lift pot, move can pot) versus baselines that train one model per task. GE-Act beats π0 and GO-1 on three of four tasks; only slightly behind on lift pot, attributed to task interference from joint training.
Long-horizon memory-intensive demo (Figure 2): Candy-packing task on Agilex Cobot Magic with only 1 hour of teleoperated data. The robot folds a deformable box, places a target inside, closes the lid (rendering the object invisible), then selects and applies the correct stamp from internal memory — demonstrating cross-step memory integration learned from GE-Base pretraining.
Task: "grasping a red cylinder from the table and placing it into a paper cup with fixed positions" on AgiBot G1. 305 demos, 40,000 training steps. The two binary axes are VidAW (initialization from GE-Base) and VidAda (task-specific video adaptation in Section 3.2). Numbers are shown with/without robot state ("S"):
| VidAW | VidAda | E2E w/ S | E2E w/o S | SR w/ S | SR w/o S |
|---|---|---|---|---|---|
| ✗ | ✗ | 0.15 | 0.30 | 0.05 | 0.11 |
| ✗ | ✓ | 0 | 0.05 | 0 | 0 |
| ✓ | ✗ | 0.81 | 0.49 | 0.64 | 0.26 |
| ✓ | ✓ | 0.89 | 0.37 | 0.76 | 0.37 |
Key takeaways: (i) training from scratch or adapting LTX-Video directly yields near-zero success — embodied pretraining is essential; (ii) GE-Base pretraining gives 64% SR / 81% E2E, rising to 76% / 89% with task-specific video adaptation; (iii) without GE-Base initialization, robot state input causes shortcut learning and reduces performance.
| Variant | SR ↑ | E2E ↑ |
|---|---|---|
| Last-layer only | 0.64 | 0.81 |
| Multi-scale fusion (ours) | 0.76 | 0.89 |
Block-wise alignment with multi-scale fusion gives +12 SR / +8 E2E over the standard "use only last visual layer" recipe used by prior VLAs.
- Data coverage and source diversity. Pretraining only on AgiBot-World-Beta (real, dual-arm, multi-view); no web video, broader robot platforms, or simulation. Few-shot transfer to Dual Franka, Agilex Cobot Magic, and RoboTwin works, but systematic OOD-embodiment robustness is underexplored.
- Embodiment scope and dexterity. Only upper-body tabletop manipulation with parallel-jaw grippers. No dexterous in-hand manipulation, tight-tolerance tool use, or whole-body navigation+manipulation.
- Real-time video modeling vs control bandwidth. Even with 5 Hz video / 30 Hz action async, highly dynamic tasks may need faster updates.
- Compute efficiency and accessibility. GE-Base + GE-Act pretraining demands many GPUs. Inference on a single commodity GPU is fine, but training cost limits adoption — they suggest distillation and quantization for the future.
Among the strongest "video-generative model is the policy backbone" bets in 2026. The paper's taxonomy of how prior work couples world models to action is worth quoting:
- (A) Video-as-policy backbones (Vid2World, VPP, UVA, Vid-MAN): serial video → action, lossy compression.
- (B) Unified video-action generators (GR-2, Cheang et al., 2024; Zhu et al. 2025 "Unified World Models"): joint generation, but pixel-level cost at inference.
- (C) WM-as-intermediate reference (Dreamitate): heuristic post-processing of generated videos.
GE sits between (A) and (B): block-wise parallel cross-attention to multi-scale latents, no recompression, slow-fast async. Where Cosmos Policy and Ctrl-World keep the world model as a substrate, Genie Envisioner unifies world-model and policy in one stack.
Versus 2026 contemporaries:
- vs π0 / π0.5 / π0.6 / π0.7: π-family policies use VLM features directly; GE-Act conditions on video-diffusion latents that explicitly encode future dynamics. The paper notes that mixing continuous diffusion losses with language objectives tends to destabilize VLM-based pipelines.
- vs GR00T N1: GR00T trains generalist humanoid manipulation via large action data; GE-Act pretrains on predictive video then ports to action with much less per-task data (1 hour for cross-embodiment transfer).
- vs OpenVLA: OpenVLA and similar VLA-style models compress visuals to language tokens; GE-Base preserves explicit pixel-level future prediction.
- vs WMPO: WMPO uses a world model for policy optimization via imagined rollouts; GE uses the world model directly as the policy substrate.
The key takeaway for the field: multi-scale, block-aligned video diffusion latents, plus slow-fast async inference, can deliver real-time control while retaining the spatial detail that language-tokenized VLAs throw away.
- OpenReview: https://openreview.net/forum?id=fHLtSxDFKC
- PDF: https://openreview.net/pdf?id=fHLtSxDFKC
- Project page: https://genie-envisioner.github.io
← Back to ICLR-2026