ICML 2026 BehaviorVLA - Heungwoo/research GitHub Wiki

From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model — a Mamba-based behavioral manifold for robust VLA control

Venue: ICML 2026 (Oral) Category: VLA Architecture Affiliations: Bing Hu, Zaijing Li, Rui Shao, Junda Chen, April Hua Liu, Wei-Shi Zheng, Liqiang Nie

Motivation, architecture, and performance of BehaviorVLA (Figure 1 from Hu et al., 2026)

Problem

VLA models degrade sharply under distribution shift — especially in sim-to-real transfer, where variations in object material, scene clutter, or camera viewpoint can cause catastrophic failure. The authors trace this brittleness to the manifold hypothesis: high-dimensional visuomotor trajectories concentrate near a low-dimensional manifold, yet standard VLAs learn mappings directly in the ambient high-dimensional space without explicit manifold constraints, so predicted actions drift off the valid task manifold under shift. Prior latent-action approaches (BeT, VQ-BeT, ACT) suffer from two limitations: (i) short-horizon temporal fragmentation — slicing trajectories into independent chunks/discrete codes loses long-term dependencies; and (ii) static execution-alignment — decoding from static latents ignores real-time execution progress, causing temporal drift.

Method

BehaviorVLA learns a temporally coherent behavioral representation via two symmetric components. It is built on the π0.5 backbone.

BehaviorVLA architecture: VBE encoder and PBD decoder (Figure 2 from Hu et al., 2026)

  • Visuomotor Behavior Encoder (VBE). A causal three-stream Mamba architecture (vision, action, and behavior streams) aggregates long-horizon trajectory information. The selective state-space model uses an input-dependent timescale Δₜ so the VBE acts as a selective information bottleneck, suppressing background clutter while preserving task events at linear O(L) complexity. After per-stream temporal filtering, mutual cross-attention fuses vision and action streams, and the behavior stream queries the joint context to distill the global task topology.
  • Manifold coordinate parameterization. The encoded trajectory is decoupled into a time-invariant global prototype z_proto (task topology, retrieved from an offline Behavior Memory Bank) and a time-variant phase state z_phase (execution progress).
  • Phase-conditioned Behavior Decoder (PBD). Operates as a Predictor–Corrector: it first unfolds a stable behavior skeleton via phase-guided topology unfolding, then refines it with geometry-guided flow matching, dynamically aligning task-level priors with real-time progress.

Training is two-phase: Phase 1 behavior-manifold learning, then Phase 2 prior-guided policy tuning (30k steps, 8×A800, batch 256, lr 5×10⁻⁵).

Results

  • RoboTwin 2.0 (Hard setting, 20 tasks × 100 rollouts): 58% average success, +37.7% over RDT.
  • LIBERO (Spatial/Object/Goal/Long, 500 rollouts/suite): 98% average, SOTA across all suites, surpassing π0.5 and OpenVLA-OFT; largest gain on LIBERO-Long (+2.2%).
  • CALVIN (ABC→D): 4.36 average rollout length.
  • Real-world (GALAXEA R1 Lite, 14-DoF bimanual): BehaviorVLA matches OpenVLA-OFT using only 50% of the demonstration data, evidencing strong data efficiency and generalization.
  • Ablation (LIBERO + Real-World): adding both VBE and PBD lifts real-world Generalization to 70.0 and Long-horizon to 55.0, versus 57.0 / 41.0 for the backbone alone.

Significance

By casting robust manipulation as learning on a low-dimensional behavioral manifold — and explicitly factorizing global task topology (prototype) from local execution progress (phase) — BehaviorVLA addresses both temporal fragmentation and static alignment in one framework. Its linear-complexity Mamba encoder scales to long-horizon demonstrations, and the 50%-data sim-to-real parity is a strong data-efficiency result for an oral-track VLA architecture paper.

Links

← Back to ICML-2026