ICML 2026 HALO - Heungwoo/research GitHub Wiki

HALO — A Unified VLA Model for Embodied Multimodal Chain-of-Thought Reasoning

Venue: ICML 2026 (Poster) Category: VLA Architecture Affiliations: HKUST, EPFL, Sun Yat-sen University (Quanxin Shou, Fangqi Zhu, et al.; corresp. Song Guo) Traction (2026-06): 2 citations (arXiv)

HALO's unified Mixture-of-Transformers architecture with three specialized experts — Multimodal Understanding, Visual Generation, and Action Prediction — sharing a self-attention mechanism (Figure 1 from Shou et al., 2026)

Problem

Vision-Language-Action (VLA) models perform well on manipulation but struggle in long-horizon or out-of-distribution scenarios because they lack explicit mechanisms for multimodal reasoning and for anticipating how the world will evolve under action. Recent work adds either textual chain-of-thought or visual subgoal prediction, but no method offers a unified human-like reasoning framework that jointly does textual reasoning, visual foresight, and action prediction.

Method

HALO is a unified VLA model that performs embodied multimodal chain-of-thought (EM-CoT) reasoning as a sequential process: textual task reasoning → visual subgoal prediction (fine-grained guidance) → EM-CoT-augmented action prediction.

  • Unified architecture. Inspired by BAGEL's harmonization of multimodal tasks, HALO uses a Mixture-of-Transformers (MoT) with three specialized experts — Multimodal Understanding, Visual Generation, and Action Prediction — that keep independent parameter sets but interact through a shared self-attention mechanism, enabling rich cross-modal collaboration while a switching mechanism controls the active modality.
  • EM-CoT data pipeline. An automatic three-phase pipeline synthesizes EM-CoT training data at scale: (1) translate continuous low-level actions into high-level motion primitives via rule-based matching, (2) use a large VLM (Qwen3-VL) to add dense textual reasoning — task narratives and subtask decomposition, and (3) designate each subtask's terminal frame as its visual subgoal, giving sparse supervision that lowers learning difficulty.
  • Two-stage training recipe. Stage 1 — Versatile Pre-training over a heterogeneous mix of VQA, Visual Generation, and Action Prediction data to build a generalist foundation. Stage 2 — EM-CoT-Augmented Fine-tuning on the synthesized reasoning/subgoal-aligned data to elicit structured multimodal reasoning.

EM-CoT data pipeline: low-level actions to motion primitives, VLM-augmented dense textual reasoning, and terminal-frame visual subgoals (Figure 2 from Shou et al., 2026)

Results

Evaluated on RoboTwin 2.0 (50 tasks) against π₀, RDT, and Diffusion Policy, with baseline numbers taken from the official RoboTwin 2.0 leaderboard.

  • Simulation. HALO reaches 80.46% on Easy and 26.44% on Hard tasks, surpassing π₀ by 34.1% and 10.1% respectively. The relative gaps over π₀ (73.5% Easy, 62.0% Hard) are large especially on Hard tasks, indicating stronger OOD robustness.
  • Ablations. Even without the explicit reasoning chain, HALO-w/o-EM-CoT already beats the strongest baseline (π₀) by +28.92 points; all training-recipe and EM-CoT components further improve success rate.
  • Real-world. Across four basic tasks (tool-use sweeping, bimanual cup nesting, inter-arm screwdriver handover, drawer placement) plus a generalization setting with visual distractions, lighting/background changes, and novel objects, HALO shows strong generalization under aggressive unseen randomization, with accurate textual reasoning and subgoal image generation.

Significance

HALO unifies textual reasoning, visual foresight, and action into a single MoT-based VLA, rather than bolting one reasoning modality onto a policy. The shared-attention three-expert design plus an automated EM-CoT data pipeline lets the model "think in words, imagine in pixels, then act" — yielding both higher success rates and notably better out-of-distribution robustness on a demanding bimanual benchmark.

Links

← Back to ICML-2026