Review HALO - Heungwoo/research GitHub Wiki
Model: HALO — unified three-expert Mixture-of-Transformers VLA (think → imagine → act) · HKUST · EPFL · Sun Yat-sen University (Quanxin Shou, Fangqi Zhu, …; corresp. Song Guo) Paper: arXiv 2602.21157 · ICML 2026 (poster, venue page) · survey entry: ICML-2026-HALO One-liner: the clearest academic instance of the three-expert MoT with a dedicated visual tower — and the one that ablates why the vision tower earns its place.
Part of the WAM+VLA hybrid family: VLA Hybrid Architectures (comparison + design guide) · siblings Motus · BagelVLA · BAGEL · DYNA-2.
-
Three experts, one shared self-attention. HALO is a Mixture-of-Transformers with a Multimodal-Understanding expert (autoregressive text), a Visual-Generation expert (diffusion / flow-matching → subgoal image), and an Action expert (flow-matching → action chunk). The experts keep independent parameters but share the self-attention; special tokens (
<visual_start>,<action_start>) route the active modality. -
Embodied Multimodal Chain-of-Thought (EM-CoT). It reasons like a person: textual reasoning
r→ visual subgoalô→ actiona, each conditioned on the last (Eqs. 1–3). "Think in words, imagine in pixels, then act." - Small backbone, big gains. Each expert is a Qwen2.5-1.5B (~4.5B total). On RoboTwin 2.0 (50 tasks): 80.5% Easy / 26.4% Hard, beating π0 by +34.1 / +10.1 points; on real Cobot Mobile ALOHA it leads π0/π0.5 on all four tasks.
- The ablations are the contribution. Removing the visual-generation data alone drops Hard success >50%; no pre-training → 0% on Hard. Both the textual chain and the visual subgoal are needed for the OOD robustness.
- It isolates the value of the third (vision) tower. Where Motus and DYNA-2 argue for a vision tower at scale, HALO ablates it at small scale: dropping visual-generation data costs >50% on Hard tasks, and the visual-subgoal branch is separately necessary. This is the cleanest published evidence that a dedicated visual-foresight tower does real work, not just regularization.
- EM-CoT unifies two previously separate CoT lines. Prior work adds either textual chain-of-thought or visual subgoal prediction; HALO fuses both into one sequential reasoning process inside one model — a structured answer to long-horizon / OOD manipulation.
- It runs the vision tower in-path (as a subgoal), yet stays a small model. Unlike full-video WAMs (3–4× latency), HALO's vision tower emits a single subgoal image, keeping the "imagine" step affordable — the mid-point of the granularity axis in Review-VLA-Hybrid-Architectures §4(a).

flowchart LR
L[instruction + obs history] --> U[Understanding expert · AR<br/>Qwen2.5-1.5B · ViT+SigLIP2/NaViT]
U -->|reasoning r| V[Visual-Generation expert · diffusion<br/>FLUX VAE, subgoal image ô]
V -->|subgoal ô| A[Action expert · flow-matching<br/>action chunk a]
U <-. shared self-attention .-> V
V <-. shared self-attention .-> A
U <-. shared self-attention .-> A
A ==> OUT[action]
- Understanding expert (AR). Qwen2.5-1.5B (28 layers, 12 heads, hidden 1536). Vision for understanding: ViT + SigLIP2 (384²→980²) with NaViT for native aspect ratios.
- Visual-Generation expert (diffusion). Predicts the subgoal image via flow-matching (MSE); pixels through a frozen FLUX VAE (8× downsample, 16 latent channels).
- Action expert (flow-matching, L₁). Low-dim continuous actions via linear projection.
- Fusion. Independent parameter sets per expert, shared self-attention; a switching mechanism (special tokens) sets the active modality. Attention masking: causal for AR text, bidirectional within a frame / causal across frames, noise tokens masked from clean context.
-
EM-CoT (Eqs. 1–3):
r ~ P(·|l,o)→ô_{t+h} ~ P(·|l,o,r)→a_{t:t+m} ~ π(·|l,o,r,ô).
Training — two stages.
-
Stage 1 — Versatile Pre-training (90k steps): VQA (LLaVA-NeXT-779k, CE) + Visual Generation (OXE + SSv2 video, flow-MSE) + Action (OXE, L₁). Loss
L = 0.25·L_CE + 0.5·L_MSE + L_L1. -
Stage 2 — EM-CoT-Augmented Fine-tuning (110k sim / 80k real): an automated 3-phase EM-CoT pipeline — (1) low-level actions → motion primitives by rule-matching, (2) Qwen3-VL adds dense textual reasoning + subtask decomposition, (3) each subtask's terminal frame = visual subgoal. Corpus: 2,500 sim demos + 320 real demos, co-trained with general VQA to prevent forgetting. Loss
L = L_r + L_ô + L_a.

RoboTwin 2.0 (50 tasks, 100 trials each; baselines from the official leaderboard):
| Setting | HALO | π0 | best baseline (RDT-1B) |
|---|---|---|---|
| Easy | 80.5% | 46.4% (+34.1) | 34.5% |
| Hard | 26.4% | 16.3% (+10.1) | 13.7% |
Hard-task highlights: Stack-Blocks-Three 37% (π0 0%); Shake-Bottle 73% (π0 60%).
Real world — Cobot Mobile ALOHA, 4 tasks × 50 trials:
| Task | HALO | π0 | π0.5 |
|---|---|---|---|
| Sweep buttons | 98% | 84% | 82% |
| Bimanual cup nesting | 92% | 64% | 68% |
| Screwdriver handover | 88% | 72% | 76% |
| Lemon → drawer | 94% | 70% | 74% |
Ablations (the core evidence):
| Pre-training config | Easy | Hard |
|---|---|---|
| Full (Vision + Text + Action) | 75.3% | 21.2% |
| w/o Visual-generation data | 58.2% | 10.5% |
| w/o Vision + Text VQA | 42.9% | 3.9% |
| w/o pre-training | 32.4% | 0% |
EM-CoT: w/o textual reasoning 77.8 / 18.3; w/o visual subgoal 76.1 / 22.5; both (HALO) 80.5 / 26.4. Even HALO-w/o-EM-CoT beats π0 by +28.9 on Easy.
Significance. HALO is the strongest controlled case that the three-expert (vision-as-tower) MoT is more than a parameter dump: each tower's data measurably drives success, EM-CoT compounds them, and the win is largest on Hard / OOD tasks — exactly where monolithic VLAs fail. It also shows the recipe works at ~4.5B, not just 8–14B.
Limitations.
- arXiv/ICML-poster; unreplicated externally. Numbers are the authors' own (leaderboard-anchored baselines).
- Vision tower runs in-path. A subgoal image is cheaper than a video rollout but still an extra generative step per chunk; no latency table is reported.
- Heavy reliance on pre-training diversity (0% Hard without it) — the recipe is data-recipe-sensitive, not just architecture-driven.
- Small backbone caps language breadth (Qwen2.5-1.5B experts); the EM-CoT text is task-scoped, not open-domain reasoning.
- No explicit limitations section; the failure envelope is inferred from ablations.
- Paper: arXiv 2602.21157 · ICML 2026: poster · survey entry: ICML-2026-HALO
- Family: VLA Hybrid Architectures · Motus · BagelVLA · BAGEL (base recipe) · DYNA-2 · Being-H0.7
- Taxonomy: VLA Architectures §4.2b · VLM↔Action Connection · World Models