Review HALO - Heungwoo/research GitHub Wiki

In-Depth Review — HALO: a three-expert VLA for Embodied Multimodal Chain-of-Thought

Model: HALO — unified three-expert Mixture-of-Transformers VLA (think → imagine → act) · HKUST · EPFL · Sun Yat-sen University (Quanxin Shou, Fangqi Zhu, …; corresp. Song Guo) Paper: arXiv 2602.21157 · ICML 2026 (poster, venue page) · survey entry: ICML-2026-HALO One-liner: the clearest academic instance of the three-expert MoT with a dedicated visual tower — and the one that ablates why the vision tower earns its place.

Part of the WAM+VLA hybrid family: VLA Hybrid Architectures (comparison + design guide) · siblings Motus · BagelVLA · BAGEL · DYNA-2.


1. TL;DR

  1. Three experts, one shared self-attention. HALO is a Mixture-of-Transformers with a Multimodal-Understanding expert (autoregressive text), a Visual-Generation expert (diffusion / flow-matching → subgoal image), and an Action expert (flow-matching → action chunk). The experts keep independent parameters but share the self-attention; special tokens (<visual_start>, <action_start>) route the active modality.
  2. Embodied Multimodal Chain-of-Thought (EM-CoT). It reasons like a person: textual reasoning r → visual subgoal ô → action a, each conditioned on the last (Eqs. 1–3). "Think in words, imagine in pixels, then act."
  3. Small backbone, big gains. Each expert is a Qwen2.5-1.5B (~4.5B total). On RoboTwin 2.0 (50 tasks): 80.5% Easy / 26.4% Hard, beating π0 by +34.1 / +10.1 points; on real Cobot Mobile ALOHA it leads π0/π0.5 on all four tasks.
  4. The ablations are the contribution. Removing the visual-generation data alone drops Hard success >50%; no pre-training → 0% on Hard. Both the textual chain and the visual subgoal are needed for the OOD robustness.

2. Why it matters

  • It isolates the value of the third (vision) tower. Where Motus and DYNA-2 argue for a vision tower at scale, HALO ablates it at small scale: dropping visual-generation data costs >50% on Hard tasks, and the visual-subgoal branch is separately necessary. This is the cleanest published evidence that a dedicated visual-foresight tower does real work, not just regularization.
  • EM-CoT unifies two previously separate CoT lines. Prior work adds either textual chain-of-thought or visual subgoal prediction; HALO fuses both into one sequential reasoning process inside one model — a structured answer to long-horizon / OOD manipulation.
  • It runs the vision tower in-path (as a subgoal), yet stays a small model. Unlike full-video WAMs (3–4× latency), HALO's vision tower emits a single subgoal image, keeping the "imagine" step affordable — the mid-point of the granularity axis in Review-VLA-Hybrid-Architectures §4(a).

3. Architecture

HALO's unified Mixture-of-Transformers — three experts (Multimodal Understanding · Visual Generation · Action Prediction) with independent parameters sharing one self-attention (Figure 1 from Shou et al., 2026, © the authors)

flowchart LR
  L[instruction + obs history] --> U[Understanding expert · AR<br/>Qwen2.5-1.5B · ViT+SigLIP2/NaViT]
  U -->|reasoning r| V[Visual-Generation expert · diffusion<br/>FLUX VAE, subgoal image ô]
  V -->|subgoal ô| A[Action expert · flow-matching<br/>action chunk a]
  U <-. shared self-attention .-> V
  V <-. shared self-attention .-> A
  U <-. shared self-attention .-> A
  A ==> OUT[action]
Loading
  • Understanding expert (AR). Qwen2.5-1.5B (28 layers, 12 heads, hidden 1536). Vision for understanding: ViT + SigLIP2 (384²→980²) with NaViT for native aspect ratios.
  • Visual-Generation expert (diffusion). Predicts the subgoal image via flow-matching (MSE); pixels through a frozen FLUX VAE (8× downsample, 16 latent channels).
  • Action expert (flow-matching, L₁). Low-dim continuous actions via linear projection.
  • Fusion. Independent parameter sets per expert, shared self-attention; a switching mechanism (special tokens) sets the active modality. Attention masking: causal for AR text, bidirectional within a frame / causal across frames, noise tokens masked from clean context.
  • EM-CoT (Eqs. 1–3): r ~ P(·|l,o) → ô_{t+h} ~ P(·|l,o,r) → a_{t:t+m} ~ π(·|l,o,r,ô).

Training — two stages.

  • Stage 1 — Versatile Pre-training (90k steps): VQA (LLaVA-NeXT-779k, CE) + Visual Generation (OXE + SSv2 video, flow-MSE) + Action (OXE, L₁). Loss L = 0.25·L_CE + 0.5·L_MSE + L_L1.
  • Stage 2 — EM-CoT-Augmented Fine-tuning (110k sim / 80k real): an automated 3-phase EM-CoT pipeline — (1) low-level actions → motion primitives by rule-matching, (2) Qwen3-VL adds dense textual reasoning + subtask decomposition, (3) each subtask's terminal frame = visual subgoal. Corpus: 2,500 sim demos + 320 real demos, co-trained with general VQA to prevent forgetting. Loss L = L_r + L_ô + L_a.

HALO's automated EM-CoT data pipeline — low-level actions → motion primitives (rule-based), VLM-augmented dense textual reasoning, and terminal-frame visual subgoals (Figure 2 from Shou et al., 2026, © the authors)


4. Results (paper-reported)

RoboTwin 2.0 (50 tasks, 100 trials each; baselines from the official leaderboard):

Setting HALO π0 best baseline (RDT-1B)
Easy 80.5% 46.4% (+34.1) 34.5%
Hard 26.4% 16.3% (+10.1) 13.7%

Hard-task highlights: Stack-Blocks-Three 37% (π0 0%); Shake-Bottle 73% (π0 60%).

Real world — Cobot Mobile ALOHA, 4 tasks × 50 trials:

Task HALO π0 π0.5
Sweep buttons 98% 84% 82%
Bimanual cup nesting 92% 64% 68%
Screwdriver handover 88% 72% 76%
Lemon → drawer 94% 70% 74%

Ablations (the core evidence):

Pre-training config Easy Hard
Full (Vision + Text + Action) 75.3% 21.2%
w/o Visual-generation data 58.2% 10.5%
w/o Vision + Text VQA 42.9% 3.9%
w/o pre-training 32.4% 0%

EM-CoT: w/o textual reasoning 77.8 / 18.3; w/o visual subgoal 76.1 / 22.5; both (HALO) 80.5 / 26.4. Even HALO-w/o-EM-CoT beats π0 by +28.9 on Easy.


5. Significance & limitations

Significance. HALO is the strongest controlled case that the three-expert (vision-as-tower) MoT is more than a parameter dump: each tower's data measurably drives success, EM-CoT compounds them, and the win is largest on Hard / OOD tasks — exactly where monolithic VLAs fail. It also shows the recipe works at ~4.5B, not just 8–14B.

Limitations.

  1. arXiv/ICML-poster; unreplicated externally. Numbers are the authors' own (leaderboard-anchored baselines).
  2. Vision tower runs in-path. A subgoal image is cheaper than a video rollout but still an extra generative step per chunk; no latency table is reported.
  3. Heavy reliance on pre-training diversity (0% Hard without it) — the recipe is data-recipe-sensitive, not just architecture-driven.
  4. Small backbone caps language breadth (Qwen2.5-1.5B experts); the EM-CoT text is task-scoped, not open-domain reasoning.
  5. No explicit limitations section; the failure envelope is inferred from ablations.

6. Links

← Back to Reviews · ICML-2026 · Home

⚠️ **GitHub.com Fallback** ⚠️