Review BAGEL - Heungwoo/research GitHub Wiki

In-Depth Review — BAGEL: the unified multimodal MoT recipe robot VLAs inherit

Model: BAGEL-7B-MoT — open unified multimodal understanding+generation model · ByteDance-Seed Paper: "Emerging Properties in Unified Multimodal Pretraining" — arXiv 2505.14683 · HF ByteDance-Seed/BAGEL-7B-MoT · blog Why it's in this wiki: BAGEL is not a robot model — but its Mixture-of-Transformers recipe is the template that BagelVLA, HALO, and the whole three-expert vision-tower family build on. This page documents the base recipe so the robot pages can point to one source.

Family: VLA Hybrid Architectures · BagelVLA · HALO · Motus.


1. TL;DR

  1. One model, understanding + generation, via Mixture-of-Transformer-Experts. BAGEL runs two experts — multimodal understanding and visual generation — with separate QKV projectors and FFNs per expert but shared attention layers. 7B active / 14B total, initialized from Qwen2.5-7B-Instruct.
  2. Dual visual encoders — the key idea robot VLAs copy. A VAE supplies pixel-level features (for generation) and a ViT supplies semantic features (for understanding), so one model both sees and draws without the two objectives fighting.
  3. Trillions of interleaved tokens. Pre-trained on interleaved text · image · video · web data; capability appears in stages (understanding+generation early → basic editing → complex editing).
  4. Emergent "world-modeling." Beyond image gen/edit/understand, it shows future-frame prediction, free-form manipulation, multiview synthesis, and world navigation — the very capabilities a WAM needs, which is why robot hybrids adopt its backbone.

2. Why it matters for VLA

  • It is the base recipe of the three-expert MoT. HALO states it is "inspired by BAGEL's harmonization of multimodal tasks"; BagelVLA literally initializes from BAGEL and adds a third (action) expert. The pattern — per-expert QKV/FFN + shared self-attention, plus VAE+ViT dual vision encoders — is what lets robot models bolt on a video/action tower without wrecking language grounding (Review-VLA-Hybrid-Architectures §2, §5).
  • It proves the fusion is interference-free at scale. BAGEL beats strong specialist models on both understanding and generation simultaneously — evidence that separate-QKV/FFN-shared-attention is a real solution to the multi-objective conflict, not a compromise.
  • Its emergent world-modeling is the bridge to WAMs. Future-frame prediction and world navigation emerging from pure multimodal pretraining is exactly the prior a WAM wants — so a BAGEL-initialized policy starts with a usable world model.

3. Architecture

BAGEL's Mixture-of-Transformer-Experts — an Understanding expert (Next-Token Prediction) and a Generation expert (Velocity Prediction), each with its own QKV + FFN but sharing one "Multi-modal Self-Attention"; separate Text Tokenizer + Und Encoder (semantic) and Gen Encoder (pixel/VAE) feed the two experts (architecture figure from arXiv 2505.14683, © the authors)

flowchart LR
  subgraph ENC[Dual visual encoders]
    VAE[VAE · pixel-level] 
    VIT[ViT · semantic-level]
  end
  IN[text · image · video · web] --> VAE
  IN --> VIT
  VAE --> G[Generation expert<br/>separate QKV/FFN]
  VIT --> U[Understanding expert<br/>separate QKV/FFN]
  U <-. shared attention layers .-> G
  U --> OU[text / understanding]
  G --> OG[image / edit / future frame]
Loading
  • MoT-Experts: understanding + generation, separate QKV + FFN per expert, shared attention. Init from Qwen2.5-7B-Instruct; 7B active / 14B total.
  • Dual vision path: VAE (pixel detail → generation) + ViT (semantics → understanding).
  • Objective: a "Next Group of Token Prediction" paradigm over interleaved multimodal tokens.
  • Data: trillions of interleaved text/image/video/web tokens across pretrain → continued training → SFT.

4. Results (paper-reported)

Task BAGEL Competitor
MME (understanding) 2388 Qwen2.5-VL 2347
MMBench 85.0 Qwen2.5-VL 83.5
GenEval (text→image) 0.88 SD3-Medium 0.74
GEdit-Bench (SC, editing) 7.36 Step1X-Edit 7.09
  • Emergent staging: understanding + generation appear early; basic editing next; complex/intelligent editing later — and free-form manipulation, multiview synthesis, and world navigation ("world-modeling" tasks) emerge with scale.

5. Significance & limitations

Significance. BAGEL is the load-bearing base of the three-expert hybrid family — the demonstration that a shared-attention, per-expert-QKV/FFN MoT with VAE+ViT dual encoders can unify understanding and generation at frontier quality. Robot VLAs (BagelVLA, HALO) inherit exactly this and add an action expert.

Limitations (from a VLA standpoint).

  1. Not a robot model. No action head, no embodiment — its world-modeling is emergent and qualitative, not benchmarked for control.
  2. Heavy (14B total). The base ticket before an action tower is even added.
  3. World-navigation / future-frame are demonstrations, not controllable dynamics — a policy still needs an action expert + robot data (what BagelVLA/HALO supply).
  4. General-purpose objective ≠ manipulation prior — the gap between "can imagine plausible frames" and "predicts contact-accurate dynamics" is the same one all pixel WAMs face (Review-World-Models §6).

6. Links

← Back to Reviews · Home

⚠️ **GitHub.com Fallback** ⚠️