Review BAGEL - Heungwoo/research GitHub Wiki
Model: BAGEL-7B-MoT — open unified multimodal understanding+generation model · ByteDance-Seed Paper: "Emerging Properties in Unified Multimodal Pretraining" — arXiv 2505.14683 · HF ByteDance-Seed/BAGEL-7B-MoT · blog Why it's in this wiki: BAGEL is not a robot model — but its Mixture-of-Transformers recipe is the template that BagelVLA, HALO, and the whole three-expert vision-tower family build on. This page documents the base recipe so the robot pages can point to one source.
Family: VLA Hybrid Architectures · BagelVLA · HALO · Motus.
- One model, understanding + generation, via Mixture-of-Transformer-Experts. BAGEL runs two experts — multimodal understanding and visual generation — with separate QKV projectors and FFNs per expert but shared attention layers. 7B active / 14B total, initialized from Qwen2.5-7B-Instruct.
- Dual visual encoders — the key idea robot VLAs copy. A VAE supplies pixel-level features (for generation) and a ViT supplies semantic features (for understanding), so one model both sees and draws without the two objectives fighting.
- Trillions of interleaved tokens. Pre-trained on interleaved text · image · video · web data; capability appears in stages (understanding+generation early → basic editing → complex editing).
- Emergent "world-modeling." Beyond image gen/edit/understand, it shows future-frame prediction, free-form manipulation, multiview synthesis, and world navigation — the very capabilities a WAM needs, which is why robot hybrids adopt its backbone.
- It is the base recipe of the three-expert MoT. HALO states it is "inspired by BAGEL's harmonization of multimodal tasks"; BagelVLA literally initializes from BAGEL and adds a third (action) expert. The pattern — per-expert QKV/FFN + shared self-attention, plus VAE+ViT dual vision encoders — is what lets robot models bolt on a video/action tower without wrecking language grounding (Review-VLA-Hybrid-Architectures §2, §5).
- It proves the fusion is interference-free at scale. BAGEL beats strong specialist models on both understanding and generation simultaneously — evidence that separate-QKV/FFN-shared-attention is a real solution to the multi-objective conflict, not a compromise.
- Its emergent world-modeling is the bridge to WAMs. Future-frame prediction and world navigation emerging from pure multimodal pretraining is exactly the prior a WAM wants — so a BAGEL-initialized policy starts with a usable world model.

flowchart LR
subgraph ENC[Dual visual encoders]
VAE[VAE · pixel-level]
VIT[ViT · semantic-level]
end
IN[text · image · video · web] --> VAE
IN --> VIT
VAE --> G[Generation expert<br/>separate QKV/FFN]
VIT --> U[Understanding expert<br/>separate QKV/FFN]
U <-. shared attention layers .-> G
U --> OU[text / understanding]
G --> OG[image / edit / future frame]
- MoT-Experts: understanding + generation, separate QKV + FFN per expert, shared attention. Init from Qwen2.5-7B-Instruct; 7B active / 14B total.
- Dual vision path: VAE (pixel detail → generation) + ViT (semantics → understanding).
- Objective: a "Next Group of Token Prediction" paradigm over interleaved multimodal tokens.
- Data: trillions of interleaved text/image/video/web tokens across pretrain → continued training → SFT.
| Task | BAGEL | Competitor |
|---|---|---|
| MME (understanding) | 2388 | Qwen2.5-VL 2347 |
| MMBench | 85.0 | Qwen2.5-VL 83.5 |
| GenEval (text→image) | 0.88 | SD3-Medium 0.74 |
| GEdit-Bench (SC, editing) | 7.36 | Step1X-Edit 7.09 |
- Emergent staging: understanding + generation appear early; basic editing next; complex/intelligent editing later — and free-form manipulation, multiview synthesis, and world navigation ("world-modeling" tasks) emerge with scale.
Significance. BAGEL is the load-bearing base of the three-expert hybrid family — the demonstration that a shared-attention, per-expert-QKV/FFN MoT with VAE+ViT dual encoders can unify understanding and generation at frontier quality. Robot VLAs (BagelVLA, HALO) inherit exactly this and add an action expert.
Limitations (from a VLA standpoint).
- Not a robot model. No action head, no embodiment — its world-modeling is emergent and qualitative, not benchmarked for control.
- Heavy (14B total). The base ticket before an action tower is even added.
- World-navigation / future-frame are demonstrations, not controllable dynamics — a policy still needs an action expert + robot data (what BagelVLA/HALO supply).
- General-purpose objective ≠ manipulation prior — the gap between "can imagine plausible frames" and "predicts contact-accurate dynamics" is the same one all pixel WAMs face (Review-World-Models §6).
- Paper: arXiv 2505.14683 · weights: HF · blog: seed.bytedance.com
- Robot descendants: BagelVLA (adds action expert) · HALO (three-expert VLA) · Motus
- Family & theory: VLA Hybrid Architectures · VLA Architectures §4.2b · World Models