ICLR 2026 villa X - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 Category: VLA Architecture — Latent action Trend tag: Latent actions · cross-embodiment · video pretraining Affiliation: Microsoft Research · Tsinghua University · Wuhan University · HKUST · Nanjing University
flowchart LR
subgraph LAM[Latent Action Model]
O1[o_t] --> IDM[IDM]
O2[o_t+K] --> IDM
IDM --> VQ[VQ codebook<br/>size 32]
VQ --> z[Latent action z_t]
z --> FDM[Visual FDM]
FDM --> Vrec[ô_t+K image]
z --> pFDM[proprio-FDM]
Q[q_t] --> pFDM
CE[embodiment context c_e<br/>dataset ID + freq] --> pFDM
pFDM --> Q2[q̂_t+1..t+K]
pFDM --> A2[â_t..t+K-1]
end
subgraph ACT[ACTor module — joint diffusion]
VLM[PaliGemma 3B VLM] --> ACTL[ACT-latent expert<br/>~300M params]
ACTL --> ACTR[ACT-robot expert<br/>~300M params]
QT[proprio q_t] --> ACTR
CE2[embodiment ctx c_e] --> ACTR
Wrist[Wrist camera ResNet-18] --> ACTR
end
z -.supervises latent token prediction.-> ACTL
Prior latent-action VLAs (LAPA, Moto-GPT, IGOR, GO-1, GR00T) compress motion between consecutive frames into a discrete codebook from visual signal alone. This works for large pixel changes but ignores motions that are critical yet visually subtle — end-effector rotations, gripper open/close. The resulting latents are physically ungrounded, and the ways prior work injects them into VLA training (LAPA: pretrained-init only; GO-1: autoregressive teacher-forcing; GR00T: latent-as-embodiment) leave information on the table. villa-X attacks both: better latent learning and better integration.
Standard recipe is z_t = IDM(o_t, o_{t+K}), ô_{t+K} = FDM(o_t, z_t), trained on visual reconstruction. villa-X adds a proprioceptive Forward Dynamics Model (proprio-FDM):
(q̂_{t+1},…,q̂_{t+K}, â_{t+1},…,â_{t+K}) = proprio-FDM(q_t, z_t, c_e)
with embodiment context c_e = (dataset ID, control frequency), embedded via learnable embeddings and concatenated with the robot state before being fed into the proprio-FDM. This disambiguates heterogeneous robot platforms so the latent itself does not have to encode "which robot." Joint loss = visual reconstruction + proprioceptive prediction + VQ commitment. For human video that has no proprio labels, only the visual term is active. The LAM uses a vector-quantization module with a codebook of size 32; the resulting latent action z_t is what downstream policy training predicts (via flow matching) and conditions on.
The policy factorizes:
π(a_{t:t+m-1}, z^K_{t:t+(n-1)K} | o_t, l, q_t, c_e) = π_robot(a | z, o, l, q_t, c_e) · π_latent(z | o, l)
Realized as 3 experts under blockwise causal attention:
- VLM (PaliGemma 3B, 224×224 images, 128-token text) — produces high-level features.
- ACT-latent — 18-layer transformer, hidden dim 1024, 8 heads, ~300M params; predicts latent action sequence (n=6).
- ACT-robot — same architecture, ~300M params; predicts robot action chunk (m=4) conditioned on VLM features, predicted latents, proprio q_t, c_e, optional wrist features.
For grouped variable x_t = (z, a) and conditioning O_t = (o_t, l, q_t, c_e):
L_τ(θ) = E[ ‖ v^θ_τ(x^τ_t, O_t) − u(x^τ_t | x_t) ‖² ]
where x^τ_t = τ x_t + (1−τ)ε and u(x^τ_t|x_t) = ε − x_t. Both the ACT-latent and ACT-robot experts are trained with flow matching, with the robot-action diffusion process conditioned (via attention) on the latent-action diffusion process so that information transfers from latent plan to robot action. (The specific τ-sampling distribution per expert is not stated in the paper's main text.)
To stop ACT-robot from short-circuiting through latent tokens, the authors apply two complementary dropout schemes:
- 50% of training steps: all robot-to-latent attention is masked.
- Otherwise: 50% of latent tokens are randomly masked.
Plus 50% wrist-camera dropout since not every dataset has wrist views.
Per-embodiment state-projection and action-projection layers (HPT design) wrap a shared transformer. Wrist camera is encoded by a ResNet-18 and fused via cross-attention into 16 tokens.
- LAM pretraining: batch 512, lr 1.5e-4, 2k linear warmup, ~4 days on 128 A100 GPUs.
- ACT pretraining (joint latent + robot): lr 5e-5, 200-step warmup, grad clip 1.0, ~4 days on 64 A100 GPUs.
- Embodiment-specific fine-tuning on each downstream platform.
- 1.6 M robot trajectories / 223.5 M frames from OpenX + AgiBot World Beta (key shares: AgiBot 20%, RT-1 9.7%, Bridge 5.47%, BC-Z 3.47%, DROID 3.46%, Kuka 1.97%, Stanford Hydra 1.61%).
- 3.6 M human-video clips: Ego4D 21.46%, EPIC-KITCHENS 6.95%, Something-Something V2 6.82%, RH20T 5.56%, HoloAssist 4.77%, HOI4D 1.99%, EgoPAT3D 0.94%, EGTEA Gaze+ 0.89%, HO-Cap 0.63%.
3-layer MLP probes trained on frozen latent actions to predict LIBERO ground-truth robot actions (eight dims: 3 pos + 4 rot + 1 gripper). Maximum-L1 error histograms show the w/pp (with proprio-FDM) variant produces strictly more low-error samples than wo/pp (visual-only).
Google robot / WidowX averages, success-rate %:
| Model | Google avg | WidowX avg |
|---|---|---|
| RT-1-X* | 49.4 | 1.1 |
| Octo-base* | 14.6 | 16.0 |
| OpenVLA* | 32.7 | 1.0 |
| RoboVLMs* | 55.3 | 13.5 |
| RoboVLMs (post-trained) | 60.8 | 37.5 |
| π0 | 58.7 | 27.1 |
| π0-FAST | 61.9 | 32.1 |
| OpenVLA-OFT | 63.0 | N/A |
| GR00T-N1.5 | 57.9 | 62.0 |
| TraceVLA | 57.3 | 27.7 |
| Magma | 62.3 | 44.8 |
| MoTo | 59.2 | N/A |
| LAPA | N/A | 57.3 |
| villa-X w/o latent (ablation) | 36.5 | 49.0 |
| villa-X (Ours) | 77.7 | 62.5 |
(* = evaluated directly after pretraining; all other baselines are evaluated after post-training. Baseline scores are cited from the original publications or related literature, with missing entries N/A.)
Per-task on Google robot (3 tasks): villa-X 98.7 / 75.0 / 59.3 (Pick / Move / Drawer → avg 77.7); on WidowX (4 tasks) 46.3 / 64.6 / 77.9 / 61.3 (Carrot / Eggplant / Spoon / Cube → avg 62.5).
LAM-design + integration ablation (Table 1 of paper; Google avg over 3 tasks, WidowX over 4):
| Latent design | Google avg | WidowX avg |
|---|---|---|
| Ours (w/pp) | 58.5 | 40.8 |
| wo/pp (visual-FDM only) | 57.4 | 32.3 |
| wo/LAM (no latent) | 35.0 | 33.1 |
| LAPA-style integration | 43.8 | 1.0 |
| GO-1-style integration | 32.8 | 14.8 |
(These are at smaller pretraining scale: 10% Fractal + 10% Bridge V2 + 100% SSv2.)
| Method | Pick-in | Pick-out | Push | Stack | Unstack | Color-OOD | Table-OOD |
|---|---|---|---|---|---|---|---|
| GR00T | 30 | 70 | 10 | 10 | 60 | 50 | 30 |
| Ours w/o latent | 40 | 80 | 30 | 60 | 70 | 40 | 30 |
| Ours | 30 | 100 | 50 | 50 | 100 | 60 | 60 |
(Table 4 of the paper compares only GR00T, Ours w/o latent, and Ours on Realman.)
Fine-tuned on 375 teleop trajectories (75 per task).
On XHand (4,000 trajectories, 13 categories — no dexterous data in pretraining):
| Task | seen | unseen |
|---|---|---|
| Pick & Place — GR-1 / GR00T / Ours w/o lat / Ours | 56 / 44 / 72 / 84 | 40 / 28 / 60 / 68 |
| Stack Cube — GR-1 / GR00T / Ours w/o lat / Ours | 15 / 20 / 70 / 75 | 5 / 0 / 40 / 50 |
| Place Cup Upright | 0 / 20 / 40 / 60 | 0 / 0 / 30 / 30 |
| Pour Water | 0 / 0 / 40 / 60 | 0 / 0 / 10 / 30 |
| Flick Ball | 40 / 30 / 50 / 50 | 10 / 0 / 30 / 40 |
Demonstrates embodiment transfer to a 12-DoF hand the model has never seen.
On a Realman arm (unseen embodiment) with symbol cards (e.g. "touch the corn"): ACT-latent rolls out latent action plans and a separately trained world model renders them. Authors qualitatively show the rendered trajectories follow open-vocabulary symbolic instructions, evidencing both embodiment-agnostic and open-vocab generalization.
- Proprio-FDM (w/pp vs wo/pp): +1.1 pp Google (58.5 vs 57.4) / +8.5 pp WidowX (40.8 vs 32.3) SIMPLER avg (Table 1).
- Latent expert (Ours vs Ours w/o latent): +41.2 pp Google (77.7 vs 36.5) / +13.5 pp WidowX (62.5 vs 49.0) in SIMPLER (Table 2).
- Embodiment context c_e (Sec. D.3): "Ours w/o context" produces higher visual-FDM and proprio-FDM reconstruction loss on validation; on the Realman novel-embodiment probing experiment, removing c_e degrades latent-quality probing (numbers in Appendix D.3).
- LAPA-style vs GO-1-style integration: both significantly underperform villa-X's joint-diffusion integration with the same data and backbone (Table 1).
- Stochastic latent-attention masking: "We found this design crucial in practice" — the masking is what prevents ACT-robot from learning trivial shortcuts through latents. (No explicit numerical sweep but the authors flag it as essential.)
- The latent expert's planning capacity is "not fully explored." For instance, sampling multiple latent plans and rejecting ones a VLM critic deems instruction-violating could improve robustness; this is left for future work.
- The framework is described as generic; richer structural cues (end-effector keypoints, human hand pose) could replace proprioception for grounding but are not investigated.
- Cost: LAM pretrain takes 4 days × 128 A100 = ~12,288 A100-hours; ACT pretrain 4 days × 64 A100 ≈ 6,144 A100-hours. Reproducibility at smaller compute is not characterized.
- vs LAPA (Ye et al. 2024): LAPA learns latents from video, then discards the latent prediction head and continues training on robot data with a new action head. villa-X jointly models latent + robot via flow-matching, never throws latents away — and outperforms the LAPA-style integration by +14.7 pp on Google SIMPLER and +39.8 pp on WidowX in the like-for-like Table 1 comparison (58.5/40.8 vs 43.8/1.0).
- vs GO-1: GO-1 autoregresses discrete latents and conditions actions on them with teacher forcing; villa-X uses joint diffusion (no teacher-forcing inconsistency between train and test). +25.7 pp Google; +26.0 pp WidowX in the same-budget Table 1 comparison (58.5/40.8 vs 32.8/14.8).
- vs GR00T-N1.5: the paper's own GR00T-N1.5 run scores 57.9/62.0 on SIMPLER (Google/WidowX, Table 2), versus villa-X 77.7/62.5; on the Realman gripper and XHand dexterous setups villa-X also outperforms GR00T on most tasks.
- vs π0 lineage: the π-series use VLM + flow-matching expert directly without a latent intermediate. In the paper's Table 2 they are competitive baselines (π0 58.7/27.1, π0-FAST 61.9/32.1, OpenVLA-OFT 63.0/N/A), all surpassed by villa-X (77.7/62.5). villa-X argues a mid-level latent plan distilled from human + robot video adds significant generalization, especially on dexterous-hand (XHand: villa-X vs GR00T +40 pp Pick&Place seen / +40 pp unseen). (Note: π0, π0-FAST, OpenVLA-OFT, TraceVLA, and Magma ARE all reported baselines in Table 2; the wiki previously stated the opposite, which was incorrect.)
- Cross-embodiment claim: zero-shot to Realman + zero-shot to 12-DoF XHand (no dexterous data in pretraining) is the strongest evidence so far that latent actions can transfer across kinematic morphologies.
- OpenReview: https://openreview.net/forum?id=y5CaJb17Fn
- Human Video Pretraining
- EgoDex
- Actions as Language
- OneTwoVLA (also π0-based, different bet on what to add)
← Back to ICLR-2026