Review VLA Hybrid Architectures - Heungwoo/research GitHub Wiki
In-Depth Review β VLA Hybrid Architectures: the three-expert Mixture-of-Transformers (vision as a separate tower)
Scope: the 2026 convergence where a VLA and a world/video model fuse into one policy β with a focused comparison of the papers that split vision (video / visual-foresight) into its own expert tower in a three-expert Mixture-of-Transformers (MoT). Companions: VLA Architectures Β§4.2b (the taxonomy this expands) Β· World Models Β· NVIDIA WAM thesis Β· WAM vs VLA Robustness. Written as a design reference for building a new architecture β Β§6 is a decision guide.
A VLA grounds language β action directly; a world/action model (WAM) inserts an imagine stage (predict how the scene evolves, then act). Through 2026 the two converged: NVIDIA's WAM blog predicts the next generation is a WAM+VLA hybrid (Review-NVIDIA-WAM-Cosmos3 Β§2.6), and the field settled on Mixture-of-Transformers (MoT) as the vehicle. This page is the deep-dive on the sharpest version of that idea β giving vision its own tower.
Two orthogonal design axes organize the whole hybrid space (full taxonomy in Review-VLA-Architecture Β§4.2b):
- Axis 1 β fusion style: modular / bolted-on (Ο0.7, Cortex 2.0) vs unified single model (this page).
- Axis 2 β where the world/vision model sits relative to the action path: in the action path (generate/consume frames to act β heavy, ~3β4Γ latency) Β· co-training tower, dropped at inference (reactive) Β· training-time critic only (π IROS 2026 β the WM never touches deployment; it scores action chunks during offline RL post-training, then is absent β AtomVLA).
The canonical unified hybrid is a three-tower Mixture-of-Transformers: an understanding/VLM tower, a vision tower (video generation or visual-foresight), and an action tower. Each tower keeps its own weights and its natural generative objective β autoregressive text for reasoning, diffusion / flow-matching for visual and action β while the towers exchange information through shared self-attention (the BAGEL recipe, Β§3: separate QKV projectors + FFNs per expert, shared attention layers).
flowchart LR
IN[obs Β· language] --> U[Understanding / VLM tower<br/>AR text Β· reasoning, subtask]
IN --> V[Vision tower<br/>diffusion β video / keyframe / latent future]
IN --> A[Action tower<br/>flow-matching β action chunk]
U <-. shared self-attention .-> V
V <-. shared self-attention .-> A
U <-. shared self-attention .-> A
A ==> OUT[action]
V -. run in-path / cheap 1-step / dropped .-> A
Why split vision off? A single VLM fine-tuned to also generate video and actions suffers gradient conflict; per-expert QKV/FFN with shared attention lets the visual-generation objective and the action objective coexist without one clobbering the other, while still cross-conditioning. It also makes the vision tower droppable at inference (Axis 2) β the key to buying a world prior without its latency.
| Model | The 3 experts (towers) | Vision tower predicts | Vision tower at inference | Scale | Open | Headline result |
|---|---|---|---|---|---|---|
| Motus (THU-ML, 2512.13030) | understanding Β· video-generation Β· action | full future video (scheduler picks the mode) | optional β UniDiffuser scheduler switches it in/out | 8B (VGM 5.0B Β· VLM 2.13B Β· act 641M Β· und 253M) | β weights+code | RoboTwin 2.0 88.66% vs X-VLA 72.80 / Ο0.5 42.98 |
| BagelVLA (RSS'26 #83, 2602.09849) | LLM Β· generation Β· action | subtask keyframe (single image) | cheap β Residual Flow Guidance, 1-step denoise (1.2 s / 48-act chunk, 40β72 Hz) | 7B (Bagel) + 2B action | β (on Bagel) | Calvin ABC-D 4.405 (Ο0 3.648); RoboTwin 2.0 75.26% clean; real 75.5% |
| HALO (ICML'26, 2602.21157) | semantic reasoning Β· visual foresight Β· action | visual subgoal | runs (subgoal, in-path) | ~4.5B (3ΓQwen2.5-1.5B) | β | RoboTwin 2.0 Easy 80.5% (Ο0 46.4, +34.1); Hard 26.4% |
| BAGEL (base recipe) (ByteDance-Seed, 2505.14683) | understanding Β· generation (2) | image (VAE pixel + ViT semantic dual encoders) | in-path (generation) | 7B active / 14B | β | the multimodal MoT the robot models above inherit |
Structural variants (not a vision-separate three-expert MoT, but the same fusion family β useful contrasts):
| Model | Towers | Vision role | At inference | Note |
|---|---|---|---|---|
| DYNA-2 | video Β· action (2) | full video co-training | dropped β reactive | text cross-attends to video; no VLM tower |
| Being-H0.7 | understanding Β· latent-query Β· action (3) | latent future (no pixels) | dropped β reactive | posterior/prior trains a deployable prior |
| Cosmos 3 (NVIDIA) | reasoning Β· generation (2) | video + action from one gen tower | in-path | reason before generate |
| Ο-0 | latent-predictive + control | latent future | dropped β reactive | reconstruction-free (humanoid) |
| LingBot-VLA (Ant) | VL backbone Β· action (2) | vision inside the VL backbone | β | Qwen2.5-VL + action expert β the 2-tower baseline |
| AtomVLA π (IROS'26) | VLA + separate latent WM (critic) | latent future (scoring, not generating) | critic-only β used in offline RL post-training, absent at deploy | LLM subtask decomposition β WM scores chunks β offline GRPO; LIBERO 97.0% |
(a) What the vision tower predicts β granularity is the real lever.
- Full video (Motus, DYNA-2, Cosmos 3): richest physical prior, heaviest. Best for dynamics-scarce pretraining; worst for latency if run in-path.
- Keyframe / subgoal (BagelVLA, HALO, Ο0.7): a single predicted image/subgoal, aimed squarely at long-horizon, multi-stage tasks (BagelVLA's motivating example: solve an arithmetic step, then place the block). Much cheaper than a rollout.
- Latent future (Being-H0.7, Ο-0): predict a representation, never pixels β cheapest, naturally reactive, but the world prior is only as good as the latent space.
(b) Per-tower generative mechanism. The consensus (explicit in HALO) is to keep each expert's natural objective: AR for text reasoning, diffusion/flow-matching for visual and action. The MoT's separate-QKV/FFN-shared-attention wiring (BAGEL) is what makes heterogeneous objectives cohabit one model.
(c) Does the vision tower run at inference? β the latency frontier. This single choice dominates deployability:
- In-path (Cosmos-Policy/DreamZero, Cosmos 3): generate/consume frames to act β strongest prior, 3β4Γ latency (Review-WAM-vs-VLA-Robustness).
- Cheap single-step (BagelVLA's RFG): one denoising step extracts predictive visual features β "foresight without full image synthesis" β recovering real-time rates (40β72 Hz).
- Dropped β reactive (DYNA-2, Being-H0.7, Ο-0): the vision tower is training-only; inference is a plain fast policy. The world prior is baked into shared weights.
- Critic-only, at post-training (π AtomVLA, IROS 2026): the WM is used neither in-path nor as a co-training tower β it scores candidate action chunks against LLM-derived subtasks in latent space during offline GRPO, enabling RL-quality post-training without online robot rollouts, then plays no role at deployment. This is the cheapest way to inject a world prior β it never costs inference latency and never needs the WM at train-time-in-the-loop; the tradeoff is the prior only shapes the policy indirectly (via reward), not the representation.
(d) How the vision tower yields action supervision from label-free video. A recurring sub-problem, solved three ways: optical-flow "delta action" (Motus), hand-pose pseudo-actions (DYNA-2), future-informed latent posterior (Being-H0.7). This is what lets the vision tower pretrain on human/web video the action tower can't.
(e) Scale, openness, evidence. Motus (8B, open) and BagelVLA (9B, open on Bagel) are the reproducible references; HALO reports only a relative +34.1% (ICML); the strongest industrial datapoints (Cosmos 3, DYNA-2) are vendor-reported and closed. No study yet ablates the third tower's marginal value against a 2-tower VL+action baseline at matched compute.
(f) IROS 2026 β what changed (π). The insight is reinforced, not overturned: convergence continues and the latent / low-inference-cost corner keeps winning. Two refinements: (1) AtomVLA adds the critic-only hybrid mode above β a third answer on Axis 2 that pushes "cheapest world prior" to its limit (no inference cost, no in-the-loop WM at train time). (2) The WAM cluster broadened β cross-embodiment world models for dexterous manipulation, DreamMimic (humanoid loco-manip via a WM), RoboDream (WM as a data factory) β confirming World Models's "one WM, many roles" thesis inside the hybrid frame (backbone Β· co-training tower Β· reactive prior Β· offline critic Β· data engine). The marginal-value gap still stands β even with AtomVLA, no head-to-head isolates the world-model contribution at matched compute.
Pros
- Interference-free multi-objective training β per-expert QKV/FFN means video-generation gradients don't corrupt the VLM's language grounding or the action head (the failure mode of monolithic VLA+video fine-tuning).
- Modular, staged training β Motus's 3-phase pipeline and BagelVLA's "add a 2B action expert last" both exploit tower separation to stage capabilities and freeze what's done.
- Droppable at inference β the vision tower can be omitted (reactive) or run cheaply (RFG), decoupling the world-prior benefit from latency β the property monolithic designs can't offer.
- Natural-objective per modality β AR text + diffusion visual/action, each at its best.
Cons
- Parameter/compute cost β 8β14B is the entry ticket; prices out small labs.
-
Loss balancing is delicate β "video training dilutes action-learning gradients" (DYNA-2, Motus); the video/action loss weight
Ξ»is a live knob. - Latency if run in-path β full generation is 3β4Γ a VLA; only cheap-1-step or dropped variants hit real-time.
- Unproven marginal value β no controlled head-to-head shows the third tower beats a well-tuned 2-tower VL+action at equal budget.
- How many towers? 2 (VL+action, or video+action) is simpler and often enough; add a third vision tower only if you need both language grounding and generative visual foresight (long-horizon, multi-stage, or human-video pretraining).
- Pick the vision granularity to your task horizon. Long-horizon/compositional β keyframe/subgoal (BagelVLA/HALO). Dynamics-scarce pretraining β full video (Motus). Latency-critical β latent (Being-H0.7/Ο-0).
- Decide the inference contract first. In-path (accept 3β4Γ latency), cheap 1-step (RFG), or dropped/reactive (co-train the vision tower, don't run it). This choice constrains everything else.
- Plan action-from-video supervision if pretraining on label-free video: optical-flow latent, hand-pose, or a posterior/prior latent.
- Keep each expert's natural objective (AR text, diffusion/flow visual+action) and fuse via separate QKV/FFN + shared self-attention (BAGEL) to avoid interference.
- Budget the Ξ» (videoβaction loss weight) and expect to tune it β it is the reported failure knob.
- Does the third (vision) tower earn its parameters vs a 2-tower VL+action at matched compute? No paper isolates this.
- Which vision granularity wins per task class (video vs keyframe vs latent)? Motus/BagelVLA/Being-H0.7 each argue a different point on the curve; no shared benchmark compares them.
- Where should the world model sit β in-path, reactive co-training tower, or offline critic? DYNA-2 and Being-H0.7 bet reactive; Cosmos 3 keeps it in-path; AtomVLA (IROS 2026) bets critic-only (WM shapes the policy via offline-RL reward, never at inference). Which mode wins on the hardest contact-rich / long-horizon tasks is open β and it may be task-dependent (in-path foresight for precise contact, critic-only for cheap long-horizon shaping).
- Is the MoT a transient stage that collapses into the dual-system (Cat F) taxonomy, or a durable family? (See Review-VLA-Architecture Β§8.)
- Core papers (in-depth): Motus (2512.13030) Β· BagelVLA (2602.09849) Β· HALO (2602.21157) Β· BAGEL (2505.14683)
- Variants: DYNA-2 Β· Being-H0.7 Β· Ο-0 Β· Cosmos 3 / NVIDIA WAM Β· Cortex 2.0 Β· AtomVLA (critic-only, IROS'26)
- IROS 2026 context: IROS 2026 survey Β§5.1
- Taxonomy & theory: VLA Architectures Β§4.2b Β· World Models Β· WAM vs VLA Robustness
- Latest Papers Β· Reviews