Review VLA Hybrid Architectures - Heungwoo/research GitHub Wiki

In-Depth Review β€” VLA Hybrid Architectures: the three-expert Mixture-of-Transformers (vision as a separate tower)

Scope: the 2026 convergence where a VLA and a world/video model fuse into one policy β€” with a focused comparison of the papers that split vision (video / visual-foresight) into its own expert tower in a three-expert Mixture-of-Transformers (MoT). Companions: VLA Architectures Β§4.2b (the taxonomy this expands) Β· World Models Β· NVIDIA WAM thesis Β· WAM vs VLA Robustness. Written as a design reference for building a new architecture β€” Β§6 is a decision guide.


1. What "VLA hybrid" means

A VLA grounds language β†’ action directly; a world/action model (WAM) inserts an imagine stage (predict how the scene evolves, then act). Through 2026 the two converged: NVIDIA's WAM blog predicts the next generation is a WAM+VLA hybrid (Review-NVIDIA-WAM-Cosmos3 Β§2.6), and the field settled on Mixture-of-Transformers (MoT) as the vehicle. This page is the deep-dive on the sharpest version of that idea β€” giving vision its own tower.

Two orthogonal design axes organize the whole hybrid space (full taxonomy in Review-VLA-Architecture Β§4.2b):

  • Axis 1 β€” fusion style: modular / bolted-on (Ο€0.7, Cortex 2.0) vs unified single model (this page).
  • Axis 2 β€” where the world/vision model sits relative to the action path: in the action path (generate/consume frames to act β€” heavy, ~3–4Γ— latency) Β· co-training tower, dropped at inference (reactive) Β· training-time critic only (πŸ†• IROS 2026 β€” the WM never touches deployment; it scores action chunks during offline RL post-training, then is absent β€” AtomVLA).

2. The three-expert MoT β€” anatomy

The canonical unified hybrid is a three-tower Mixture-of-Transformers: an understanding/VLM tower, a vision tower (video generation or visual-foresight), and an action tower. Each tower keeps its own weights and its natural generative objective β€” autoregressive text for reasoning, diffusion / flow-matching for visual and action β€” while the towers exchange information through shared self-attention (the BAGEL recipe, Β§3: separate QKV projectors + FFNs per expert, shared attention layers).

flowchart LR
  IN[obs Β· language] --> U[Understanding / VLM tower<br/>AR text Β· reasoning, subtask]
  IN --> V[Vision tower<br/>diffusion β€” video / keyframe / latent future]
  IN --> A[Action tower<br/>flow-matching β€” action chunk]
  U <-. shared self-attention .-> V
  V <-. shared self-attention .-> A
  U <-. shared self-attention .-> A
  A ==> OUT[action]
  V -. run in-path / cheap 1-step / dropped .-> A
Loading

Why split vision off? A single VLM fine-tuned to also generate video and actions suffers gradient conflict; per-expert QKV/FFN with shared attention lets the visual-generation objective and the action objective coexist without one clobbering the other, while still cross-conditioning. It also makes the vision tower droppable at inference (Axis 2) β€” the key to buying a world prior without its latency.


3. The papers β€” three-expert, vision as a separate tower

Model The 3 experts (towers) Vision tower predicts Vision tower at inference Scale Open Headline result
Motus (THU-ML, 2512.13030) understanding Β· video-generation Β· action full future video (scheduler picks the mode) optional β€” UniDiffuser scheduler switches it in/out 8B (VGM 5.0B Β· VLM 2.13B Β· act 641M Β· und 253M) βœ… weights+code RoboTwin 2.0 88.66% vs X-VLA 72.80 / Ο€0.5 42.98
BagelVLA (RSS'26 #83, 2602.09849) LLM Β· generation Β· action subtask keyframe (single image) cheap β€” Residual Flow Guidance, 1-step denoise (1.2 s / 48-act chunk, 40–72 Hz) 7B (Bagel) + 2B action βœ… (on Bagel) Calvin ABC-D 4.405 (Ο€0 3.648); RoboTwin 2.0 75.26% clean; real 75.5%
HALO (ICML'26, 2602.21157) semantic reasoning Β· visual foresight Β· action visual subgoal runs (subgoal, in-path) ~4.5B (3Γ—Qwen2.5-1.5B) β€” RoboTwin 2.0 Easy 80.5% (Ο€0 46.4, +34.1); Hard 26.4%
BAGEL (base recipe) (ByteDance-Seed, 2505.14683) understanding Β· generation (2) image (VAE pixel + ViT semantic dual encoders) in-path (generation) 7B active / 14B βœ… the multimodal MoT the robot models above inherit

Structural variants (not a vision-separate three-expert MoT, but the same fusion family β€” useful contrasts):

Model Towers Vision role At inference Note
DYNA-2 video Β· action (2) full video co-training dropped β†’ reactive text cross-attends to video; no VLM tower
Being-H0.7 understanding Β· latent-query Β· action (3) latent future (no pixels) dropped β†’ reactive posterior/prior trains a deployable prior
Cosmos 3 (NVIDIA) reasoning Β· generation (2) video + action from one gen tower in-path reason before generate
Ο‰-0 latent-predictive + control latent future dropped β†’ reactive reconstruction-free (humanoid)
LingBot-VLA (Ant) VL backbone Β· action (2) vision inside the VL backbone β€” Qwen2.5-VL + action expert β€” the 2-tower baseline
AtomVLA πŸ†• (IROS'26) VLA + separate latent WM (critic) latent future (scoring, not generating) critic-only β€” used in offline RL post-training, absent at deploy LLM subtask decomposition β†’ WM scores chunks β†’ offline GRPO; LIBERO 97.0%

4. Comparative analysis (the axes that matter)

(a) What the vision tower predicts β€” granularity is the real lever.

  • Full video (Motus, DYNA-2, Cosmos 3): richest physical prior, heaviest. Best for dynamics-scarce pretraining; worst for latency if run in-path.
  • Keyframe / subgoal (BagelVLA, HALO, Ο€0.7): a single predicted image/subgoal, aimed squarely at long-horizon, multi-stage tasks (BagelVLA's motivating example: solve an arithmetic step, then place the block). Much cheaper than a rollout.
  • Latent future (Being-H0.7, Ο‰-0): predict a representation, never pixels β€” cheapest, naturally reactive, but the world prior is only as good as the latent space.

(b) Per-tower generative mechanism. The consensus (explicit in HALO) is to keep each expert's natural objective: AR for text reasoning, diffusion/flow-matching for visual and action. The MoT's separate-QKV/FFN-shared-attention wiring (BAGEL) is what makes heterogeneous objectives cohabit one model.

(c) Does the vision tower run at inference? β€” the latency frontier. This single choice dominates deployability:

  • In-path (Cosmos-Policy/DreamZero, Cosmos 3): generate/consume frames to act β†’ strongest prior, 3–4Γ— latency (Review-WAM-vs-VLA-Robustness).
  • Cheap single-step (BagelVLA's RFG): one denoising step extracts predictive visual features β€” "foresight without full image synthesis" β€” recovering real-time rates (40–72 Hz).
  • Dropped β†’ reactive (DYNA-2, Being-H0.7, Ο‰-0): the vision tower is training-only; inference is a plain fast policy. The world prior is baked into shared weights.
  • Critic-only, at post-training (πŸ†• AtomVLA, IROS 2026): the WM is used neither in-path nor as a co-training tower β€” it scores candidate action chunks against LLM-derived subtasks in latent space during offline GRPO, enabling RL-quality post-training without online robot rollouts, then plays no role at deployment. This is the cheapest way to inject a world prior β€” it never costs inference latency and never needs the WM at train-time-in-the-loop; the tradeoff is the prior only shapes the policy indirectly (via reward), not the representation.

(d) How the vision tower yields action supervision from label-free video. A recurring sub-problem, solved three ways: optical-flow "delta action" (Motus), hand-pose pseudo-actions (DYNA-2), future-informed latent posterior (Being-H0.7). This is what lets the vision tower pretrain on human/web video the action tower can't.

(e) Scale, openness, evidence. Motus (8B, open) and BagelVLA (9B, open on Bagel) are the reproducible references; HALO reports only a relative +34.1% (ICML); the strongest industrial datapoints (Cosmos 3, DYNA-2) are vendor-reported and closed. No study yet ablates the third tower's marginal value against a 2-tower VL+action baseline at matched compute.

(f) IROS 2026 β€” what changed (πŸ†•). The insight is reinforced, not overturned: convergence continues and the latent / low-inference-cost corner keeps winning. Two refinements: (1) AtomVLA adds the critic-only hybrid mode above β€” a third answer on Axis 2 that pushes "cheapest world prior" to its limit (no inference cost, no in-the-loop WM at train time). (2) The WAM cluster broadened β€” cross-embodiment world models for dexterous manipulation, DreamMimic (humanoid loco-manip via a WM), RoboDream (WM as a data factory) β€” confirming World Models's "one WM, many roles" thesis inside the hybrid frame (backbone Β· co-training tower Β· reactive prior Β· offline critic Β· data engine). The marginal-value gap still stands β€” even with AtomVLA, no head-to-head isolates the world-model contribution at matched compute.


5. Why separate the vision tower β€” pros & cons

Pros

  • Interference-free multi-objective training β€” per-expert QKV/FFN means video-generation gradients don't corrupt the VLM's language grounding or the action head (the failure mode of monolithic VLA+video fine-tuning).
  • Modular, staged training β€” Motus's 3-phase pipeline and BagelVLA's "add a 2B action expert last" both exploit tower separation to stage capabilities and freeze what's done.
  • Droppable at inference β€” the vision tower can be omitted (reactive) or run cheaply (RFG), decoupling the world-prior benefit from latency β€” the property monolithic designs can't offer.
  • Natural-objective per modality β€” AR text + diffusion visual/action, each at its best.

Cons

  • Parameter/compute cost β€” 8–14B is the entry ticket; prices out small labs.
  • Loss balancing is delicate β€” "video training dilutes action-learning gradients" (DYNA-2, Motus); the video/action loss weight Ξ» is a live knob.
  • Latency if run in-path β€” full generation is 3–4Γ— a VLA; only cheap-1-step or dropped variants hit real-time.
  • Unproven marginal value β€” no controlled head-to-head shows the third tower beats a well-tuned 2-tower VL+action at equal budget.

6. Design guide β€” building a new three-expert hybrid

  1. How many towers? 2 (VL+action, or video+action) is simpler and often enough; add a third vision tower only if you need both language grounding and generative visual foresight (long-horizon, multi-stage, or human-video pretraining).
  2. Pick the vision granularity to your task horizon. Long-horizon/compositional β†’ keyframe/subgoal (BagelVLA/HALO). Dynamics-scarce pretraining β†’ full video (Motus). Latency-critical β†’ latent (Being-H0.7/Ο‰-0).
  3. Decide the inference contract first. In-path (accept 3–4Γ— latency), cheap 1-step (RFG), or dropped/reactive (co-train the vision tower, don't run it). This choice constrains everything else.
  4. Plan action-from-video supervision if pretraining on label-free video: optical-flow latent, hand-pose, or a posterior/prior latent.
  5. Keep each expert's natural objective (AR text, diffusion/flow visual+action) and fuse via separate QKV/FFN + shared self-attention (BAGEL) to avoid interference.
  6. Budget the Ξ» (video↔action loss weight) and expect to tune it β€” it is the reported failure knob.

7. Open questions

  1. Does the third (vision) tower earn its parameters vs a 2-tower VL+action at matched compute? No paper isolates this.
  2. Which vision granularity wins per task class (video vs keyframe vs latent)? Motus/BagelVLA/Being-H0.7 each argue a different point on the curve; no shared benchmark compares them.
  3. Where should the world model sit β€” in-path, reactive co-training tower, or offline critic? DYNA-2 and Being-H0.7 bet reactive; Cosmos 3 keeps it in-path; AtomVLA (IROS 2026) bets critic-only (WM shapes the policy via offline-RL reward, never at inference). Which mode wins on the hardest contact-rich / long-horizon tasks is open β€” and it may be task-dependent (in-path foresight for precise contact, critic-only for cheap long-horizon shaping).
  4. Is the MoT a transient stage that collapses into the dual-system (Cat F) taxonomy, or a durable family? (See Review-VLA-Architecture Β§8.)

8. Links

← Back to Home Β· Reviews

⚠️ **GitHub.com Fallback** ⚠️