ICLR 2026 villa X - Heungwoo/research GitHub Wiki

villa-X — Enhanced Latent Action Modeling in VLAs

Venue: ICLR 2026 Category: VLA Architecture — Latent action Trend tag: Latent actions · cross-embodiment · video pretraining Affiliation: Microsoft Research · Tsinghua University · Wuhan University · HKUST · Nanjing University

Approach diagram

flowchart LR
  subgraph LAM[Latent Action Model]
    O1[o_t] --> IDM[IDM]
    O2[o_t+K] --> IDM
    IDM --> VQ[VQ codebook<br/>size 32]
    VQ --> z[Latent action z_t]
    z --> FDM[Visual FDM]
    FDM --> Vrec[ô_t+K image]
    z --> pFDM[proprio-FDM]
    Q[q_t] --> pFDM
    CE[embodiment context c_e<br/>dataset ID + freq] --> pFDM
    pFDM --> Q2[q̂_t+1..t+K]
    pFDM --> A2[â_t..t+K-1]
  end
  subgraph ACT[ACTor module — joint diffusion]
    VLM[PaliGemma 3B VLM] --> ACTL[ACT-latent expert<br/>~300M params]
    ACTL --> ACTR[ACT-robot expert<br/>~300M params]
    QT[proprio q_t] --> ACTR
    CE2[embodiment ctx c_e] --> ACTR
    Wrist[Wrist camera ResNet-18] --> ACTR
  end
  z -.supervises latent token prediction.-> ACTL
Loading

Problem

Prior latent-action VLAs (LAPA, Moto-GPT, IGOR, GO-1, GR00T) compress motion between consecutive frames into a discrete codebook from visual signal alone. This works for large pixel changes but ignores motions that are critical yet visually subtle — end-effector rotations, gripper open/close. The resulting latents are physically ungrounded, and the ways prior work injects them into VLA training (LAPA: pretrained-init only; GO-1: autoregressive teacher-forcing; GR00T: latent-as-embodiment) leave information on the table. villa-X attacks both: better latent learning and better integration.

Detailed Method

LAM (Latent Action Model)

Standard recipe is z_t = IDM(o_t, o_{t+K}), ô_{t+K} = FDM(o_t, z_t), trained on visual reconstruction. villa-X adds a proprioceptive Forward Dynamics Model (proprio-FDM):

(q̂_{t+1},…,q̂_{t+K}, â_{t+1},…,â_{t+K}) = proprio-FDM(q_t, z_t, c_e)

with embodiment context c_e = (dataset ID, control frequency), embedded via learnable embeddings and concatenated with the robot state before being fed into the proprio-FDM. This disambiguates heterogeneous robot platforms so the latent itself does not have to encode "which robot." Joint loss = visual reconstruction + proprioceptive prediction + VQ commitment. For human video that has no proprio labels, only the visual term is active. The LAM uses a vector-quantization module with a codebook of size 32; the resulting latent action z_t is what downstream policy training predicts (via flow matching) and conditions on.

ACT (ACTor module) — joint diffusion of latent + robot actions

The policy factorizes:

π(a_{t:t+m-1}, z^K_{t:t+(n-1)K} | o_t, l, q_t, c_e) = π_robot(a | z, o, l, q_t, c_e) · π_latent(z | o, l)

Realized as 3 experts under blockwise causal attention:

  • VLM (PaliGemma 3B, 224×224 images, 128-token text) — produces high-level features.
  • ACT-latent — 18-layer transformer, hidden dim 1024, 8 heads, ~300M params; predicts latent action sequence (n=6).
  • ACT-robot — same architecture, ~300M params; predicts robot action chunk (m=4) conditioned on VLM features, predicted latents, proprio q_t, c_e, optional wrist features.

Joint flow-matching objective

For grouped variable x_t = (z, a) and conditioning O_t = (o_t, l, q_t, c_e):

L_τ(θ) = E[ ‖ v^θ_τ(x^τ_t, O_t) − u(x^τ_t | x_t) ‖² ]

where x^τ_t = τ x_t + (1−τ)ε and u(x^τ_t|x_t) = ε − x_t. Both the ACT-latent and ACT-robot experts are trained with flow matching, with the robot-action diffusion process conditioned (via attention) on the latent-action diffusion process so that information transfers from latent plan to robot action. (The specific τ-sampling distribution per expert is not stated in the paper's main text.)

Stochastic attention masking (key trick)

To stop ACT-robot from short-circuiting through latent tokens, the authors apply two complementary dropout schemes:

  1. 50% of training steps: all robot-to-latent attention is masked.
  2. Otherwise: 50% of latent tokens are randomly masked.

Plus 50% wrist-camera dropout since not every dataset has wrist views.

HPT-style policy head

Per-embodiment state-projection and action-projection layers (HPT design) wrap a shared transformer. Wrist camera is encoded by a ResNet-18 and fused via cross-attention into 16 tokens.

Three-stage training

  1. LAM pretraining: batch 512, lr 1.5e-4, 2k linear warmup, ~4 days on 128 A100 GPUs.
  2. ACT pretraining (joint latent + robot): lr 5e-5, 200-step warmup, grad clip 1.0, ~4 days on 64 A100 GPUs.
  3. Embodiment-specific fine-tuning on each downstream platform.

Pretraining data mixture (Table 5 of paper)

  • 1.6 M robot trajectories / 223.5 M frames from OpenX + AgiBot World Beta (key shares: AgiBot 20%, RT-1 9.7%, Bridge 5.47%, BC-Z 3.47%, DROID 3.46%, Kuka 1.97%, Stanford Hydra 1.61%).
  • 3.6 M human-video clips: Ego4D 21.46%, EPIC-KITCHENS 6.95%, Something-Something V2 6.82%, RH20T 5.56%, HoloAssist 4.77%, HOI4D 1.99%, EgoPAT3D 0.94%, EGTEA Gaze+ 0.89%, HO-Cap 0.63%.

Comprehensive Results

LAM-quality probing (Section 4.1)

3-layer MLP probes trained on frozen latent actions to predict LIBERO ground-truth robot actions (eight dims: 3 pos + 4 rot + 1 gripper). Maximum-L1 error histograms show the w/pp (with proprio-FDM) variant produces strictly more low-error samples than wo/pp (visual-only).

SIMPLER (Table 2: full comparison; Table 1: latent-integration ablation)

Google robot / WidowX averages, success-rate %:

Model Google avg WidowX avg
RT-1-X* 49.4 1.1
Octo-base* 14.6 16.0
OpenVLA* 32.7 1.0
RoboVLMs* 55.3 13.5
RoboVLMs (post-trained) 60.8 37.5
π0 58.7 27.1
π0-FAST 61.9 32.1
OpenVLA-OFT 63.0 N/A
GR00T-N1.5 57.9 62.0
TraceVLA 57.3 27.7
Magma 62.3 44.8
MoTo 59.2 N/A
LAPA N/A 57.3
villa-X w/o latent (ablation) 36.5 49.0
villa-X (Ours) 77.7 62.5

(* = evaluated directly after pretraining; all other baselines are evaluated after post-training. Baseline scores are cited from the original publications or related literature, with missing entries N/A.)

Per-task on Google robot (3 tasks): villa-X 98.7 / 75.0 / 59.3 (Pick / Move / Drawer → avg 77.7); on WidowX (4 tasks) 46.3 / 64.6 / 77.9 / 61.3 (Carrot / Eggplant / Spoon / Cube → avg 62.5).

LAM-design + integration ablation (Table 1 of paper; Google avg over 3 tasks, WidowX over 4):

Latent design Google avg WidowX avg
Ours (w/pp) 58.5 40.8
wo/pp (visual-FDM only) 57.4 32.3
wo/LAM (no latent) 35.0 33.1
LAPA-style integration 43.8 1.0
GO-1-style integration 32.8 14.8

(These are at smaller pretraining scale: 10% Fractal + 10% Bridge V2 + 100% SSv2.)

Real Realman gripper (Table 4; 10 trials/task)

Method Pick-in Pick-out Push Stack Unstack Color-OOD Table-OOD
GR00T 30 70 10 10 60 50 30
Ours w/o latent 40 80 30 60 70 40 30
Ours 30 100 50 50 100 60 60

(Table 4 of the paper compares only GR00T, Ours w/o latent, and Ours on Realman.)

Fine-tuned on 375 teleop trajectories (75 per task).

Real XArm + XHand 12-DoF dexterous (Table 3; 10–50 trials)

On XHand (4,000 trajectories, 13 categories — no dexterous data in pretraining):

Task seen unseen
Pick & Place — GR-1 / GR00T / Ours w/o lat / Ours 56 / 44 / 72 / 84 40 / 28 / 60 / 68
Stack Cube — GR-1 / GR00T / Ours w/o lat / Ours 15 / 20 / 70 / 75 5 / 0 / 40 / 50
Place Cup Upright 0 / 20 / 40 / 60 0 / 0 / 30 / 30
Pour Water 0 / 0 / 40 / 60 0 / 0 / 10 / 30
Flick Ball 40 / 30 / 50 / 50 10 / 0 / 30 / 40

Demonstrates embodiment transfer to a 12-DoF hand the model has never seen.

Zero-shot latent planning

On a Realman arm (unseen embodiment) with symbol cards (e.g. "touch the corn"): ACT-latent rolls out latent action plans and a separately trained world model renders them. Authors qualitatively show the rendered trajectories follow open-vocabulary symbolic instructions, evidencing both embodiment-agnostic and open-vocab generalization.

Ablation Studies

  • Proprio-FDM (w/pp vs wo/pp): +1.1 pp Google (58.5 vs 57.4) / +8.5 pp WidowX (40.8 vs 32.3) SIMPLER avg (Table 1).
  • Latent expert (Ours vs Ours w/o latent): +41.2 pp Google (77.7 vs 36.5) / +13.5 pp WidowX (62.5 vs 49.0) in SIMPLER (Table 2).
  • Embodiment context c_e (Sec. D.3): "Ours w/o context" produces higher visual-FDM and proprio-FDM reconstruction loss on validation; on the Realman novel-embodiment probing experiment, removing c_e degrades latent-quality probing (numbers in Appendix D.3).
  • LAPA-style vs GO-1-style integration: both significantly underperform villa-X's joint-diffusion integration with the same data and backbone (Table 1).
  • Stochastic latent-attention masking: "We found this design crucial in practice" — the masking is what prevents ACT-robot from learning trivial shortcuts through latents. (No explicit numerical sweep but the authors flag it as essential.)

Limitations (as stated by authors, Sec. 5)

  1. The latent expert's planning capacity is "not fully explored." For instance, sampling multiple latent plans and rejecting ones a VLM critic deems instruction-violating could improve robustness; this is left for future work.
  2. The framework is described as generic; richer structural cues (end-effector keypoints, human hand pose) could replace proprioception for grounding but are not investigated.
  3. Cost: LAM pretrain takes 4 days × 128 A100 = ~12,288 A100-hours; ACT pretrain 4 days × 64 A100 ≈ 6,144 A100-hours. Reproducibility at smaller compute is not characterized.

Significance & Positioning

  • vs LAPA (Ye et al. 2024): LAPA learns latents from video, then discards the latent prediction head and continues training on robot data with a new action head. villa-X jointly models latent + robot via flow-matching, never throws latents away — and outperforms the LAPA-style integration by +14.7 pp on Google SIMPLER and +39.8 pp on WidowX in the like-for-like Table 1 comparison (58.5/40.8 vs 43.8/1.0).
  • vs GO-1: GO-1 autoregresses discrete latents and conditions actions on them with teacher forcing; villa-X uses joint diffusion (no teacher-forcing inconsistency between train and test). +25.7 pp Google; +26.0 pp WidowX in the same-budget Table 1 comparison (58.5/40.8 vs 32.8/14.8).
  • vs GR00T-N1.5: the paper's own GR00T-N1.5 run scores 57.9/62.0 on SIMPLER (Google/WidowX, Table 2), versus villa-X 77.7/62.5; on the Realman gripper and XHand dexterous setups villa-X also outperforms GR00T on most tasks.
  • vs π0 lineage: the π-series use VLM + flow-matching expert directly without a latent intermediate. In the paper's Table 2 they are competitive baselines (π0 58.7/27.1, π0-FAST 61.9/32.1, OpenVLA-OFT 63.0/N/A), all surpassed by villa-X (77.7/62.5). villa-X argues a mid-level latent plan distilled from human + robot video adds significant generalization, especially on dexterous-hand (XHand: villa-X vs GR00T +40 pp Pick&Place seen / +40 pp unseen). (Note: π0, π0-FAST, OpenVLA-OFT, TraceVLA, and Magma ARE all reported baselines in Table 2; the wiki previously stated the opposite, which was incorrect.)
  • Cross-embodiment claim: zero-shot to Realman + zero-shot to 12-DoF XHand (no dexterous data in pretraining) is the strongest evidence so far that latent actions can transfer across kinematic morphologies.

Links

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️