ICML 2026 BehaviorVLA - Heungwoo/research GitHub Wiki
From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model — a Mamba-based behavioral manifold for robust VLA control
Venue: ICML 2026 (Oral) Category: VLA Architecture Affiliations: Bing Hu, Zaijing Li, Rui Shao, Junda Chen, April Hua Liu, Wei-Shi Zheng, Liqiang Nie

Problem
VLA models degrade sharply under distribution shift — especially in sim-to-real transfer, where variations in object material, scene clutter, or camera viewpoint can cause catastrophic failure. The authors trace this brittleness to the manifold hypothesis: high-dimensional visuomotor trajectories concentrate near a low-dimensional manifold, yet standard VLAs learn mappings directly in the ambient high-dimensional space without explicit manifold constraints, so predicted actions drift off the valid task manifold under shift. Prior latent-action approaches (BeT, VQ-BeT, ACT) suffer from two limitations: (i) short-horizon temporal fragmentation — slicing trajectories into independent chunks/discrete codes loses long-term dependencies; and (ii) static execution-alignment — decoding from static latents ignores real-time execution progress, causing temporal drift.
Method
BehaviorVLA learns a temporally coherent behavioral representation via two symmetric components. It is built on the π0.5 backbone.

- Visuomotor Behavior Encoder (VBE). A causal three-stream Mamba architecture (vision, action, and behavior streams) aggregates long-horizon trajectory information. The selective state-space model uses an input-dependent timescale Δₜ so the VBE acts as a selective information bottleneck, suppressing background clutter while preserving task events at linear O(L) complexity. After per-stream temporal filtering, mutual cross-attention fuses vision and action streams, and the behavior stream queries the joint context to distill the global task topology.
- Manifold coordinate parameterization. The encoded trajectory is decoupled into a time-invariant global prototype
z_proto(task topology, retrieved from an offline Behavior Memory Bank) and a time-variant phase statez_phase(execution progress). - Phase-conditioned Behavior Decoder (PBD). Operates as a Predictor–Corrector: it first unfolds a stable behavior skeleton via phase-guided topology unfolding, then refines it with geometry-guided flow matching, dynamically aligning task-level priors with real-time progress.
Training is two-phase: Phase 1 behavior-manifold learning, then Phase 2 prior-guided policy tuning (30k steps, 8×A800, batch 256, lr 5×10⁻⁵).
Results
- RoboTwin 2.0 (Hard setting, 20 tasks × 100 rollouts): 58% average success, +37.7% over RDT.
- LIBERO (Spatial/Object/Goal/Long, 500 rollouts/suite): 98% average, SOTA across all suites, surpassing π0.5 and OpenVLA-OFT; largest gain on LIBERO-Long (+2.2%).
- CALVIN (ABC→D): 4.36 average rollout length.
- Real-world (GALAXEA R1 Lite, 14-DoF bimanual): BehaviorVLA matches OpenVLA-OFT using only 50% of the demonstration data, evidencing strong data efficiency and generalization.
- Ablation (LIBERO + Real-World): adding both VBE and PBD lifts real-world Generalization to 70.0 and Long-horizon to 55.0, versus 57.0 / 41.0 for the backbone alone.
Significance
By casting robust manipulation as learning on a low-dimensional behavioral manifold — and explicitly factorizing global task topology (prototype) from local execution progress (phase) — BehaviorVLA addresses both temporal fragmentation and static alignment in one framework. Its linear-complexity Mamba encoder scales to long-horizon demonstrations, and the 50%-data sim-to-real parity is a strong data-efficiency result for an oral-track VLA architecture paper.
Links
- arXiv: 2605.22671
- ICML 2026: https://icml.cc/virtual/2026/poster/66596
← Back to ICML-2026