ICLR 2026 TwinVLA - Heungwoo/research GitHub Wiki

TwinVLA β€” Bimanual Manipulation by Composing Two Single-Arm VLAs

Venue: ICLR 2026 Authors: Hokyun Im (Yonsei + Microsoft Research), Euijin Jeong (Yonsei), Andrey Kolobov (Microsoft Research), Jianlong Fu (Microsoft Research), Youngwoon Lee (Yonsei) Project page: https://jellyho.github.io/TwinVLA/ Category: VLA Architecture β€” Bimanual / data efficiency Trend tag: Modular composition / data-scarce bimanual

Approach diagram

flowchart LR
  OXE[OXE single-arm data ~800h] --> Pre[SingleVLA pretrain<br/>Eagle2-1B backbone, 0.8B params<br/>5x H100, 5 days, 120k steps]
  Pre --> Dup[Duplicate VLM only<br/>share vision encoder + DiT head<br/>arm-specific proprio encoders]
  Dup --> Join[Joint Attention across two VLMs<br/>+ causal cross-arm mask]
  Dup --> Moe[MoE over shared tokens<br/>l, I_ego routed: w_left * FFN_l + 1-w_left * FFN_r]
  Join --> Twin[TwinVLA 1.3B]
  Moe --> Twin
  Re[Attention re-weighting<br/>preserve pretrained modality importance] --> Twin
  Twin --> FT[Bimanual fine-tune<br/>1x L40S, 2 days, 100k steps<br/>~50 demos per task]
  FT --> Out[Bimanual action chunk<br/>A^R_t, A^L_t via flow matching]
Loading

Problem

State-of-the-art bimanual VLAs (Ο€0 with ~10,900 hours of proprietary bimanual data; RDT-1B with ~2,400 hours of mixed pretraining and ~1,440 H100-days of compute) are unreproducible for most labs. Existing monolithic cross-embodiment approaches (Octo, NVIDIA Gr00T) handle heterogeneity via embodiment-specific decoders or zero-padded action spaces but still need substantial bimanual data. Can a strong single-arm pretrained VLA be composed with a copy of itself to handle bimanual tasks, with bimanual data used only for target-task fine-tuning?

Detailed Method

SingleVLA backbone (Appendix A). Eagle2-1B vision-language model, language head removed β†’ 0.8B params. Pretrained on ~800h subset of OpenX-Embodiment (RT-1, BridgeV2, Kuka filtered, DROID, BC-Z filtered, Stanford Hydra, etc. β€” Table 2). All actions converted to absolute end-effector pose with 6D rotation (Zhou et al. 2019) for embodiment-agnostic transfer; all datasets resampled to 20 Hz via interpolation for frequency matching (inspired by Ο€0-FAST's DCT). Action chunk size 20, sampling step 10. Action head is a DiT (Peebles & Xie 2023). Training objective is conditional flow matching (Eq. 1).

Architecture (Sec. 4).

  1. Selective duplication: duplicate only the VLM backbone; share the vision encoder and DiT action head. Each arm has its own lightweight proprioception encoder. Final model is 1.3B (vs. RDT-1B 1.2B).
  2. Joint Attention (Sec. 4.2, Algorithm 2): shared self-attention across both VLM backbones β€” concatenate Q, K, V from left + right + shared streams, attend, split back to streams. Other layers (FFN, projections) remain arm-specific. Unlike Ο€0 (which links one VLM to an action head), TwinVLA links two VLMs directly.
  3. Causal joint attention mask (Fig. 3a): lower-triangular within each arm's region; shared tokens fully accessible to both; each arm attends to half the other arm's tokens β€” symmetric cross-arm interaction without violating autoregressive constraints.
  4. MoE over shared inputs (Sec. 4.3, Eq. 3): shared (l, I_ego)_t routed via MoE(x) = w_left Β· FFN_left(x) + (1 βˆ’ w_left) Β· FFN_right(x) with w_left from a small softmax router. Other shared-token layers (Projection, LayerNorm) use output-averaging task-arithmetic. Reduces VRAM 21%, enabling batch=8 on a single 40 GB GPU.
  5. Attention re-weighting (Sec. 4.3): re-scales attention scores for the shared modality to preserve pretrained modality importance β€” reduces initial fine-tuning loss by 40%.

Hyperparameters (Table 3). SingleVLA: bs 256, lr 1e-4 cosine, AdamW, weight decay 1e-5, 120k steps. TwinVLA fine-tune: bs 8, same lr/optimizer, 100k steps, 5% warmup, vision backbone frozen, no image augmentation during fine-tune.

Comprehensive Results

Real-world (Anubis dual 6-DoF arms, parallel-jaw grippers, 1 ego + 2 wrist cams; Fig. 5, Table 6). 3 long-horizon tasks Γ— ~50 demos Γ— 20 rollouts each. Per-task final-subtask completion rate (full task success):

Task (final subtask) DP (271M) RDT-1B (1.2B) TwinVLA (1.3B) Ο€0 (3.3B, SOTA ref.)
Carrot to bag (close bag) 15.0 35.0 65.0 65.0
Brush to dustpan (put onto dustpan) 35.0 40.0 80.0 80.0
Take towel off (entirely off) 20.0 60.0 55.0 65.0

TwinVLA significantly outperforms RDT-1B and DP on average and reaches comparable performance with the Ο€0 skyline while trained only on target data. Per-task bottlenecks: carrot insertion (precise bag opening), brush insertion into dustpan, and the final towel unfolding that requires a successful arm-to-arm switch (where RDT-1B's 60% edges TwinVLA's 55%).

Simulation β€” RoboTwin 2.0 (50 tasks, Table 9) and Tabletop-Sim (4 single-tasks + language task, Table 8). RoboTwin rows are the official 50-task averages; Tabletop-Sim rows average the 4 single-tasks (Dish drainer, Handover box, Lift box, Shoes table):

Suite DP RDT-1B TwinVLA Ο€0
RoboTwin Easy 28.0 34.5 42.0 46.4
RoboTwin Hard 0.6 13.7 8.9 16.3
Tabletop-Sim Easy 24.9 61.6 75.8 72.5
Tabletop-Sim Hard 23.6 38.9 42.9 44.0

TwinVLA outperforms RDT-1B in most settings and even beats the Ο€0 skyline on Tabletop-Sim Easy (the more coordination-heavy benchmark). RoboTwin Hard is the only suite where RDT-1B clearly keeps pace (TwinVLA 8.9 vs. RDT-1B 13.7).

Data efficiency (Fig. 7). Tabletop-Sim Easy, varying demos per task: TwinVLA at 20 demos already approaches RDT-1B at 50 demos; at 50 demos, TwinVLA leads RDT-1B by a wide margin.

Compute (Fig. 2). Total: TwinVLA needs ~800h single-arm + ~50 episodes bimanual + 25 H100-days. RDT-1B: ~2,400h + ~1,440 H100-days. Ο€0: ~10,900h + >1,000 H100-days.

Robustness (Sec. 5.6). Robustness is evaluated via the Tabletop-Sim Hard setting (unseen textures, object models, distractor objects) and RoboTwin's unseen evaluation language instructions. On Tabletop-Sim Hard, TwinVLA outperforms RDT-1B by 3.3%, indicating retained robustness to unseen scenes despite no bimanual pretraining.

Language following ("Put X box into Y pot", 3 box Γ— 2 pot = 6 instruction combos; Fig. 8a, Table 8 sim column): TwinVLA 80.6 (Tabletop-Sim) beats RDT-1B (55.5) and even Ο€0 (79.2). VLAs often disregard instructions after fine-tuning; authors credit careful fine-tuning that preserves SingleVLA pretrained knowledge.

Ablation Studies

Sequential ablation (Fig. 8b) on real-world + Tabletop-Sim Easy, peeling components off the full TwinVLA:

Variant Sim Ξ”SR Real Ξ”SR Cumulative consequence
Full TwinVLA β€” β€” baseline
w/o Attention re-weighting βˆ’1.1 βˆ’1.2 +40% initial fine-tuning loss; mitigates pretrainβ†’finetune input distribution shift
w/o MoE additional βˆ’1.1 β€” +28% token length, +21% VRAM; MoE removes redundant shared-input processing
w/o Joint attention additional βˆ’4.0 additional βˆ’8.3 Largest drop β€” confirms joint attention as the critical cross-arm coordination mechanism
Scratch (no OXE pretrain) βˆ’4.6 vs. full βˆ’32.9 vs. full Single-arm pretraining is the single biggest factor in real

Twin structure vs. monolithic. TwinVLA beats RDT-1B (matched ~1.2-1.3B params) by 16.2% real, 5.0% sim, 25.1% language-following. The inductive bias of two-tower coordination outperforms monolithic learning at the same parameter budget.

Limitations stated by authors

  1. Catastrophic forgetting of single-arm skills. Because TwinVLA adapts a pretrained single-arm VLA's representations for bimanual tasks, the model forgets its single-arm manipulation skills after fine-tuning. Future work on mechanisms that prevent this forgetting could help integrate more diverse data, improve explainability, and improve unseen-task generalization.
  2. Absolute EEF action space was chosen for embodiment-agnostic transfer (joint positions are DOF-specific and unsuitable for transfer); relative or shared cross-embodiment action representations (Chi et al. 2024) are flagged as future work.

Additionally observed in experiments (not formally framed as limitations):

  • TwinVLA stays below the Ο€0 skyline on the real-world tasks overall (Ο€0 uses ~10,900h proprietary bimanual data and 3.3B params), and on the hardest precision step (Take towel "entirely off") RDT-1B's 60% even edges TwinVLA's 55%.
  • RoboTwin Hard is the only sim suite where TwinVLA loses to RDT-1B (8.9 vs. 13.7).

Significance & Positioning

TwinVLA is the most concrete "compose, don't scale" answer to the bimanual data scarcity problem. The contributions are:

  1. Architectural inductive bias (twin VLMs + joint attention + cross-arm mask) is empirically more important than scale for bimanual coordination at fixed param budget β€” RDT-1B is a controlled comparison at matched size.
  2. No bimanual pretraining required. All bimanual learning happens in the target-task fine-tune, with 50 demos.
  3. Modular reuse of existing single-arm VLAs. Authors note "existing pretrained models can also be used" for SingleVLA β€” the recipe generalizes if you have a 1B-class single-arm VLA at hand.
  4. Compute-efficient. 25 H100-days vs. RDT-1B's 1,440 puts strong bimanual VLAs within reach of single-GPU groups.

Compared to:

  • Compose Your Policies β€” also modular but at the policy-ensemble level; TwinVLA composes inside the backbone via joint attention.
  • Ο€0.6 β€” the data-scale upper bound TwinVLA narrows the gap toward.
  • X-VLA β€” cross-embodiment monolithic approach.
  • VLBiMan β€” also tackles bimanual data scarcity, but via one-shot demo + VLM grounding rather than two pretrained policies.

The neuroscience analogy (SMA + corpus callosum coordinating arm-specific motor circuits) is a useful framing rather than load-bearing argument, but the empirical result β€” that the inductive bias of two-tower coordination outperforms monolithic at matched scale and beats Ο€0 on dexterous Tabletop-Sim β€” is hard to dismiss.

Links

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️