ICLR 2026 TwinVLA - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 Authors: Hokyun Im (Yonsei + Microsoft Research), Euijin Jeong (Yonsei), Andrey Kolobov (Microsoft Research), Jianlong Fu (Microsoft Research), Youngwoon Lee (Yonsei) Project page: https://jellyho.github.io/TwinVLA/ Category: VLA Architecture β Bimanual / data efficiency Trend tag: Modular composition / data-scarce bimanual
flowchart LR
OXE[OXE single-arm data ~800h] --> Pre[SingleVLA pretrain<br/>Eagle2-1B backbone, 0.8B params<br/>5x H100, 5 days, 120k steps]
Pre --> Dup[Duplicate VLM only<br/>share vision encoder + DiT head<br/>arm-specific proprio encoders]
Dup --> Join[Joint Attention across two VLMs<br/>+ causal cross-arm mask]
Dup --> Moe[MoE over shared tokens<br/>l, I_ego routed: w_left * FFN_l + 1-w_left * FFN_r]
Join --> Twin[TwinVLA 1.3B]
Moe --> Twin
Re[Attention re-weighting<br/>preserve pretrained modality importance] --> Twin
Twin --> FT[Bimanual fine-tune<br/>1x L40S, 2 days, 100k steps<br/>~50 demos per task]
FT --> Out[Bimanual action chunk<br/>A^R_t, A^L_t via flow matching]
State-of-the-art bimanual VLAs (Ο0 with ~10,900 hours of proprietary bimanual data; RDT-1B with ~2,400 hours of mixed pretraining and ~1,440 H100-days of compute) are unreproducible for most labs. Existing monolithic cross-embodiment approaches (Octo, NVIDIA Gr00T) handle heterogeneity via embodiment-specific decoders or zero-padded action spaces but still need substantial bimanual data. Can a strong single-arm pretrained VLA be composed with a copy of itself to handle bimanual tasks, with bimanual data used only for target-task fine-tuning?
SingleVLA backbone (Appendix A). Eagle2-1B vision-language model, language head removed β 0.8B params. Pretrained on ~800h subset of OpenX-Embodiment (RT-1, BridgeV2, Kuka filtered, DROID, BC-Z filtered, Stanford Hydra, etc. β Table 2). All actions converted to absolute end-effector pose with 6D rotation (Zhou et al. 2019) for embodiment-agnostic transfer; all datasets resampled to 20 Hz via interpolation for frequency matching (inspired by Ο0-FAST's DCT). Action chunk size 20, sampling step 10. Action head is a DiT (Peebles & Xie 2023). Training objective is conditional flow matching (Eq. 1).
Architecture (Sec. 4).
- Selective duplication: duplicate only the VLM backbone; share the vision encoder and DiT action head. Each arm has its own lightweight proprioception encoder. Final model is 1.3B (vs. RDT-1B 1.2B).
- Joint Attention (Sec. 4.2, Algorithm 2): shared self-attention across both VLM backbones β concatenate Q, K, V from left + right + shared streams, attend, split back to streams. Other layers (FFN, projections) remain arm-specific. Unlike Ο0 (which links one VLM to an action head), TwinVLA links two VLMs directly.
- Causal joint attention mask (Fig. 3a): lower-triangular within each arm's region; shared tokens fully accessible to both; each arm attends to half the other arm's tokens β symmetric cross-arm interaction without violating autoregressive constraints.
-
MoE over shared inputs (Sec. 4.3, Eq. 3): shared
(l, I_ego)_trouted viaMoE(x) = w_left Β· FFN_left(x) + (1 β w_left) Β· FFN_right(x)withw_leftfrom a small softmax router. Other shared-token layers (Projection, LayerNorm) use output-averaging task-arithmetic. Reduces VRAM 21%, enabling batch=8 on a single 40 GB GPU. - Attention re-weighting (Sec. 4.3): re-scales attention scores for the shared modality to preserve pretrained modality importance β reduces initial fine-tuning loss by 40%.
Hyperparameters (Table 3). SingleVLA: bs 256, lr 1e-4 cosine, AdamW, weight decay 1e-5, 120k steps. TwinVLA fine-tune: bs 8, same lr/optimizer, 100k steps, 5% warmup, vision backbone frozen, no image augmentation during fine-tune.
Real-world (Anubis dual 6-DoF arms, parallel-jaw grippers, 1 ego + 2 wrist cams; Fig. 5, Table 6). 3 long-horizon tasks Γ ~50 demos Γ 20 rollouts each. Per-task final-subtask completion rate (full task success):
| Task (final subtask) | DP (271M) | RDT-1B (1.2B) | TwinVLA (1.3B) | Ο0 (3.3B, SOTA ref.) |
|---|---|---|---|---|
| Carrot to bag (close bag) | 15.0 | 35.0 | 65.0 | 65.0 |
| Brush to dustpan (put onto dustpan) | 35.0 | 40.0 | 80.0 | 80.0 |
| Take towel off (entirely off) | 20.0 | 60.0 | 55.0 | 65.0 |
TwinVLA significantly outperforms RDT-1B and DP on average and reaches comparable performance with the Ο0 skyline while trained only on target data. Per-task bottlenecks: carrot insertion (precise bag opening), brush insertion into dustpan, and the final towel unfolding that requires a successful arm-to-arm switch (where RDT-1B's 60% edges TwinVLA's 55%).
Simulation β RoboTwin 2.0 (50 tasks, Table 9) and Tabletop-Sim (4 single-tasks + language task, Table 8). RoboTwin rows are the official 50-task averages; Tabletop-Sim rows average the 4 single-tasks (Dish drainer, Handover box, Lift box, Shoes table):
| Suite | DP | RDT-1B | TwinVLA | Ο0 |
|---|---|---|---|---|
| RoboTwin Easy | 28.0 | 34.5 | 42.0 | 46.4 |
| RoboTwin Hard | 0.6 | 13.7 | 8.9 | 16.3 |
| Tabletop-Sim Easy | 24.9 | 61.6 | 75.8 | 72.5 |
| Tabletop-Sim Hard | 23.6 | 38.9 | 42.9 | 44.0 |
TwinVLA outperforms RDT-1B in most settings and even beats the Ο0 skyline on Tabletop-Sim Easy (the more coordination-heavy benchmark). RoboTwin Hard is the only suite where RDT-1B clearly keeps pace (TwinVLA 8.9 vs. RDT-1B 13.7).
Data efficiency (Fig. 7). Tabletop-Sim Easy, varying demos per task: TwinVLA at 20 demos already approaches RDT-1B at 50 demos; at 50 demos, TwinVLA leads RDT-1B by a wide margin.
Compute (Fig. 2). Total: TwinVLA needs ~800h single-arm + ~50 episodes bimanual + 25 H100-days. RDT-1B: ~2,400h + ~1,440 H100-days. Ο0: ~10,900h + >1,000 H100-days.
Robustness (Sec. 5.6). Robustness is evaluated via the Tabletop-Sim Hard setting (unseen textures, object models, distractor objects) and RoboTwin's unseen evaluation language instructions. On Tabletop-Sim Hard, TwinVLA outperforms RDT-1B by 3.3%, indicating retained robustness to unseen scenes despite no bimanual pretraining.
Language following ("Put X box into Y pot", 3 box Γ 2 pot = 6 instruction combos; Fig. 8a, Table 8 sim column): TwinVLA 80.6 (Tabletop-Sim) beats RDT-1B (55.5) and even Ο0 (79.2). VLAs often disregard instructions after fine-tuning; authors credit careful fine-tuning that preserves SingleVLA pretrained knowledge.
Sequential ablation (Fig. 8b) on real-world + Tabletop-Sim Easy, peeling components off the full TwinVLA:
| Variant | Sim ΞSR | Real ΞSR | Cumulative consequence |
|---|---|---|---|
| Full TwinVLA | β | β | baseline |
| w/o Attention re-weighting | β1.1 | β1.2 | +40% initial fine-tuning loss; mitigates pretrainβfinetune input distribution shift |
| w/o MoE | additional β1.1 | β | +28% token length, +21% VRAM; MoE removes redundant shared-input processing |
| w/o Joint attention | additional β4.0 | additional β8.3 | Largest drop β confirms joint attention as the critical cross-arm coordination mechanism |
| Scratch (no OXE pretrain) | β4.6 vs. full | β32.9 vs. full | Single-arm pretraining is the single biggest factor in real |
Twin structure vs. monolithic. TwinVLA beats RDT-1B (matched ~1.2-1.3B params) by 16.2% real, 5.0% sim, 25.1% language-following. The inductive bias of two-tower coordination outperforms monolithic learning at the same parameter budget.
- Catastrophic forgetting of single-arm skills. Because TwinVLA adapts a pretrained single-arm VLA's representations for bimanual tasks, the model forgets its single-arm manipulation skills after fine-tuning. Future work on mechanisms that prevent this forgetting could help integrate more diverse data, improve explainability, and improve unseen-task generalization.
- Absolute EEF action space was chosen for embodiment-agnostic transfer (joint positions are DOF-specific and unsuitable for transfer); relative or shared cross-embodiment action representations (Chi et al. 2024) are flagged as future work.
Additionally observed in experiments (not formally framed as limitations):
- TwinVLA stays below the Ο0 skyline on the real-world tasks overall (Ο0 uses ~10,900h proprietary bimanual data and 3.3B params), and on the hardest precision step (Take towel "entirely off") RDT-1B's 60% even edges TwinVLA's 55%.
- RoboTwin Hard is the only sim suite where TwinVLA loses to RDT-1B (8.9 vs. 13.7).
TwinVLA is the most concrete "compose, don't scale" answer to the bimanual data scarcity problem. The contributions are:
- Architectural inductive bias (twin VLMs + joint attention + cross-arm mask) is empirically more important than scale for bimanual coordination at fixed param budget β RDT-1B is a controlled comparison at matched size.
- No bimanual pretraining required. All bimanual learning happens in the target-task fine-tune, with 50 demos.
- Modular reuse of existing single-arm VLAs. Authors note "existing pretrained models can also be used" for SingleVLA β the recipe generalizes if you have a 1B-class single-arm VLA at hand.
- Compute-efficient. 25 H100-days vs. RDT-1B's 1,440 puts strong bimanual VLAs within reach of single-GPU groups.
Compared to:
- Compose Your Policies β also modular but at the policy-ensemble level; TwinVLA composes inside the backbone via joint attention.
- Ο0.6 β the data-scale upper bound TwinVLA narrows the gap toward.
- X-VLA β cross-embodiment monolithic approach.
- VLBiMan β also tackles bimanual data scarcity, but via one-shot demo + VLM grounding rather than two pretrained policies.
The neuroscience analogy (SMA + corpus callosum coordinating arm-specific motor circuits) is a useful framing rather than load-bearing argument, but the empirical result β that the inductive bias of two-tower coordination outperforms monolithic at matched scale and beats Ο0 on dexterous Tabletop-Sim β is hard to dismiss.
- OpenReview: https://openreview.net/forum?id=jG9W6nAwVz
- Project: https://jellyho.github.io/TwinVLA/
- Compose Your Policies β policy composition
- Ο0.6
- X-VLA
- VLBiMan β sibling bimanual data-efficient approach
- Survey: VLA & Manipulation
β Back to ICLR-2026