Review VLA Architecture Categories - Heungwoo/research GitHub Wiki

VLA Architectures โ€” Categories Aโ€“M, compared (companion to Review-VLA-Architecture)

Split out of the main VLA Architectures review for faster GitHub-wiki rendering. This is the per-category deep-dive (ยง5 of that review): pros, cons, exemplars, and when to pick each of the 13 families.

Category A โ€” Autoregressive action tokens

  • Mechanism: Extend the VLM's vocabulary with discretized EE-pose-delta bins; generate actions left-to-right with standard cross-entropy.
  • Exemplars: OpenVLA, RT-X, ฯ€0-FAST, RT-2, VLA-0.
  • Pros: โœ… Zero architectural surgery ยท โœ… inherits LLM training/inference tooling ยท โœ… trivially extensible to any output type.
  • Cons: โŒ Discretization caps precision ยท โŒ token-sequential decoding is slow ยท โŒ rigid decoding order.
  • Key datapoint: VLA-0 (NVIDIA 2025) shows that representing actions directly as text beats ฯ€0.5-KI, OpenVLA-OFT, SmolVLA, GR00T-N1, MolmoAct on LIBERO โ€” reopens the AR-vs-diffusion debate.
  • Pick when: You want the simplest possible recipe and LLM tooling compatibility.

Category B โ€” Flow-matching action expert

  • Mechanism: Separate continuous-valued head attached to the VLM; produces a full action chunk in a handful of ODE integration steps by matching a learned vector field.
  • Exemplars: ฯ€0 โ†’ ฯ€0.5 โ†’ ฯ€0.6 โ†’ ฯ€0.7 (canonical), RFS (residual), FLOWER (efficient 950M). NeurIPS 2025 additions: Knowledge Insulation (PI Spotlight โ€” the published training recipe that enables this category to coexist with a VLM), Real-Time Chunking (async chunk inpainting for boundary smoothness), ReinFlow (learnable noise injection โ†’ RL-trainable flow policies). ICRA 2026 additions: FPO (Flow Policy Optimization โ€” a likelihood-free RL fine-tuning objective that reuses the conditional-flow-matching loss differential as the PPO importance ratio, so off-the-shelf ฯ€0 is improvable from sparse reward with no sampler change; LIBERO avg 87.2), ACG (Action Coherence Guidance โ€” a flow-specific test-time guidance term that suppresses the jitter/noise-sensitivity that flow policies' high generative capacity introduces under imitation). The ICRA cohort's clear signal: flow matching is now the default continuous-action head, and its two open problems โ€” reward fine-tuning and action coherence โ€” are where the work has moved.
  • Pros: โœ… Smooth continuous actions ยท โœ… 5 Euler steps on H100 โ†’ 63 ms on ฯ€0.6 ยท โœ… clean separation of reasoning vs. control.
  • Cons: โŒ Two-stage objective is finicky; action-head gradients corrupt VLM features โ†’ Knowledge Insulation needed ยท โŒ separate head complicates the pipeline.
  • Pick when: Production system with latency budget and continuous-action smoothness requirements.

Category C โ€” Continuous diffusion action head

  • Mechanism: Same structural split as B (separate head), trained with denoising-diffusion loss rather than flow matching; multi-step sampling.
  • Exemplars: Diffusion Policy (ancestor), RDT-1B (1.2B + Physically Interpretable UAS), DexVLA (~1B plug-in), RoboDual (specialist + generalist).
  • Pros: โœ… Models multimodal action distributions well ยท โœ… mature theory ยท โœ… stable training.
  • Cons: โŒ More NFE at inference than flow matching ยท โŒ separate head complicates architecture.
  • Pick when: Training stability matters and you can afford extra denoising steps.

Category D โ€” Discrete diffusion inside the VLM

  • Mechanism: Actions are tokenized (like A) but generated via masked discrete diffusion within the same transformer โ€” adaptive easy-first unmasking, sometimes with re-masking for uncertain tokens.
  • Exemplars: Discrete Diffusion VLA (96.3% LIBERO), Unified Diffusion VLA (joint future-frame + action), dVLA (with CoT), DIVA & Fast-dVLA (latency opt), MMaDA-VLA, Dream-VLA (whole backbone is a diffusion LM).
  • Pros: โœ… Unified transformer (one objective) ยท โœ… parallel decoding ยท โœ… adaptive order ยท โœ… inherits VLM KV-cache and decoding ecosystem.
  • Cons: โŒ Discretization loss persists ยท โŒ newer โ€” training recipes less mature ยท โŒ latency still catching up to flow matching (DIVA targets parity).
  • Pick when: You want the unification argument and can tolerate early-adopter risk.

Category E โ€” World-model / video-based action

๐Ÿ“– Key empirical reference for E4 (video-backbone WAMs): WAM vs VLA Robustness โ€” Huawei + UToronto, Mar 2026 (arXiv 2603.22078). The first controlled benchmark of 7 VLAs + 2 hybrids + 4 WAMs on LIBERO-Plus + the new RoboTwin 2.0-Plus, with the headline mixed verdict: WAMs win noise/light/layout (Cosmos-Policy 82.2% LIB, LingBot-VA 74.2% RT) but lose camera + robot-state, ฯ€0.5 leads LIBERO-Plus overall (85.7%), and WAM inference latency is 4.8ร—โ€“83ร— ฯ€0.5. Empirically grounds the E4 = video-generation-backbone-as-policy sub-pattern below.

This category covers every architectural pattern where a generative model of future observations is part of the policy โ€” whether as a regularizer, a data source, an environment, a backbone, or a substitute for action generation entirely. The ยง4.2 diagram lays out the five sub-patterns; below is per-sub-pattern detail.

ICRA 2026 adds a goal-state variant. Goal-VLA (ICRA 2026) uses an image-generative VLM as an object-centric world model: it synthesizes a goal image from instruction + observation, validates it through a Reflection-through-Synthesis loop, then recovers the object's rigid transform (feature matching + point-cloud registration) and hands it to a training-free motion planner. No action-labeled data and no learned action head โ€” the generated goal state is the interface. Zero-shot, it averages 59.9% on RLBench (8 tasks) where ฯ€0 scores 0.0% and OpenVLA 0.2%; the reflection loop carries a 40.0% baseline to 88.8%. It is the world-model-as-prompt analogue to E2 (generate, then act) but skips the policy entirely, sitting at the E5/Category-F boundary (geometry-derived control via a hierarchical VLM planner).

Why this category exploded in 2025โ€“2026. Three forces converged: (1) video-generation models (LTX-Video, OpenSora, Wan2.1, Cosmos, DynamiCrafter) became cheap enough to fine-tune on robot data; (2) action-conditioned world models with bounded long-horizon error (Ctrl-World's 20+ s consistency, WMPO's reward F1 โ‰ฅ 0.95) made E3-style policy optimization viable; (3) ICLR 2026 confirmed strong empirical wins โ€” Genie-Envisioner's 1M-episode AgiBot-World-Beta scale, WMPO's 53โ†’70% real-robot lift, ViPRA's +16% SIMPLER. The thread is now the largest growth area in the ICLR 2026 VLA cohort.

E1 โ€” World model as auxiliary loss

  • Mechanism: Predict future frames / depth / segmentation as an auxiliary objective on top of VLM features; action head still drives behavior.
  • Exemplars: Cosmos Policy (NVIDIA Cosmos foundation + control tokens), Vid2World (DynamiCrafter + Diffusion Forcing + Causal Action Injection โ€” CS:GO FVD โˆ’71%, RT-1 FVD 18.5 vs 24.2 baselines), Genie Envisioner (GE-Base: LTX-Video 2B trained on 1M AgiBot episodes), Geometry-aware 4D Video (Stanford+MIT+TRI โ€” joint geometry + appearance grounding), WorldVLA, DreamVLA (multi-modal world-knowledge forecasting: depth + geom + seg + dynamic-region mask), VideoVLA.
  • Pros: โœ… Strong representation regularizer ยท โœ… no inference-time generation cost (WM is off-policy at deploy) ยท โœ… composes with any action head (A/B/C/D).
  • Cons: โŒ Loss-balancing finicky ยท โŒ doesn't directly improve action prediction โ€” gain is via shared features only.
  • Pick when: You already have a VLA that works and want extra grounding without changing the action path.

E2 โ€” World model as data factory

  • Mechanism: Train a controllable video generator on real robot data + novel language prompts; harvest the generated rollouts (with re-captioning + IDM-derived pseudo-actions) as synthetic SFT data.
  • Exemplars: DreamGen (NVIDIA's GR00T-Dreams blueprint โ€” claim: 36h sim โ‰ˆ 3 months of teleop), ViPRA (NSVQ 8-codebook + DINOv2 video pretrain โ†’ flow head; +16% SIMPLER, +13% real, 22 Hz, 240k pretrain + 50k finetune steps on 8ร—H100 over 312h), Genie Envisioner (the SFT half of the GE-Base/GE-Act pipeline), "Robotic Manipulation by Imitating Generated Videos Without Physical Demonstrations" (ICLR 2026).
  • Pros: โœ… Decouples from real data collection ยท โœ… enables novel-verb training data ยท โœ… benefits from web-video pretraining of the generator.
  • Cons: โŒ Filtering generated rollouts is hard โ€” N1 used a commercial-LLM judge on 8-frame samples ยท โŒ generator hallucinations propagate to policy ยท โŒ generator + IDM training is itself costly.
  • Pick when: You have an abundant world model + scarce target-robot data.

E3 โ€” World model as environment for policy optimization

  • Mechanism: Use the learned video / state predictor as a simulator โ€” rollout policy inside it, compute imagined rewards, optimize with policy-gradient or model-based RL. Distinct from E2 because the policy queries the WM at training time, not just learns from frozen rollouts.
  • Exemplars: Ctrl-World (pose-conditioned controllable WM, 20+ s consistency for closed-loop policy evaluation), WMPO (OpenSora world model + VideoMAE reward model F1 โ‰ฅ 0.95 โ†’ GRPO policy update; 4-task Mimicgen 33.6 โ†’ 47.1 โ†’ 57.6 across budgets; real-world 53 โ†’ 70%), WorldGym (world-model-as-benchmark), RLVR-World (RL for world-model quality, NeurIPS 2025).
  • Pros: โœ… Avoids expensive real rollouts during RL ยท โœ… closed-loop evaluation pipeline ยท โœ… scales with WM quality, not robot count.
  • Cons: โŒ Sim2real-style WM-to-reality gap reappears ยท โŒ WM artifacts can short-circuit the reward (model-exploiting policies) ยท โŒ training-time inference cost is high.
  • Pick when: You want RL on expensive embodiments and you can train a high-fidelity WM cheaply.

E4 โ€” VAM: video-generation model replaces the VLM backbone

  • Mechanism: Skip the VLM entirely. The backbone is a pretrained video-generation model (Cosmos, LTX-Video, DynamiCrafter); a small action decoder rides on top. Argument: web-video pretraining beats web-image+text pretraining for embodied tasks.
  • Exemplars: mimic-video (Dec 2025 โ€” "VAM" position piece), DiT4DiT (dual-DiT video-action, 2603.10448), S-VAM (shortcut VAM via self-distilling foresight, 2603.16195), Genie Envisioner (GE-Act: 160M action decoder block-aligned with LTX-Video 2B features, 5 Hz video + 30 Hz action async), Cosmos-Policy (ICLR-2026-Cosmos-Policy), LingBot-VA (Wan2.2-5B), Fast-WAM (Wan2.2-5B, joint-denoise + skip-video-at-test), DreamZero (Wan2.1-14B), GigaWorld-Policy (action-conditioned video).
  • Pros: โœ… Exploits true internet-scale video data ยท โœ… video features encode dynamics natively ยท โœ… "scale beats architecture" empirical bet ยท โœ… strong visual-perturbation robustness (noise/light/layout) per the Huawei robustness study โ€” Cosmos-Policy 82.2% LIBERO-Plus, LingBot-VA 74.2% RoboTwin 2.0-Plus.
  • Cons: โŒ Loses language-grounding flexibility โ€” verb generalization unclear ยท โŒ video backbones are 1โ€“3B params with high latency ยท โŒ data is even larger than VLM-style ยท โŒ weak on geometric perturbations (camera viewpoint, robot init state) โ€” video priors don't transfer to scene geometry; latency 4.8ร—โ€“83ร— ฯ€0.5 even with denoising-step reductions (Review-WAM-vs-VLA-Robustness).
  • Open empirical reference: "Do WAMs Generalize Better than VLAs?" Huawei Mar 2026 โ€” the controlled benchmark across 7 VLAs + 2 hybrids + 4 WAMs on LIBERO-Plus + RoboTwin 2.0-Plus.
  • Pick when: You believe web-video pretraining is the right prior, you have the compute, you can tolerate the latency, and visual-perturbation robustness matters more than camera-geometry robustness.

E5 โ€” Geometry-first: predict 4D / point cloud โ†’ IK

  • Mechanism: Predict future point cloud, 4D Gaussian field, or scene geometry; derive action via inverse kinematics / contact search. Skip action token generation entirely.
  • Exemplars: Avi (2510.21746, NeurIPS 2025 Workshop โ€” predict future point cloud โ†’ IK; sidesteps cross-embodiment action-vocabulary problem), Geometry-aware 4D Video (E1 + E5 hybrid โ€” predicts video and geometry for the action policy), PA3FF (part-aware 3D feature field โ€” could be either L or E5 depending on use).
  • Pros: โœ… No action-vocabulary commitment ยท โœ… naturally cross-embodiment ยท โœ… explicit physical grounding.
  • Cons: โŒ Generating future geometry is harder than future frames ยท โŒ IK adds a second optimization at inference ยท โŒ contact-rich manipulation requires physics-aware planning, not just IK.
  • Pick when: Cross-embodiment is the primary axis and contact dynamics are tractable.

When E vs. its alternatives

  • E1 vs. plain B/C/D: add E1 if you want a free representation regularizer with no inference cost. Not a substitute for the action head.
  • E2 vs. teleop scaling: E2 only wins if your WM is genuinely better than what you'd get from 100h additional teleop on the target robot.
  • E3 vs. real-robot RL: E3 is the path when real RL is impossible (expensive embodiment, safety constraints). Otherwise real RL is still cleaner.
  • E4 vs. B (ฯ€-style flow matching): E4 is the contrarian bet โ€” strong on video-rich domains, weaker on language-rich ones.
  • E5 vs. category J (sensor-augmented 3D): E5 generates geometry; J senses geometry. Different commitments.

Pick category E when: No robot data for the target (E2/E4), expensive embodiment without sim (E3), strong video-pretraining ecosystem (E1/E4), or cross-embodiment is the primary axis (E5).

Category F โ€” Hierarchical / dual-system / MoE

๐Ÿ“– Deeper treatment of the cognitive framing (Kahneman โ†’ Bengio โ†’ Chiriatti, Figure Helix-02, Sharpa CraftNet, ~26 dual/triple-system papers): see Review: System 0 / 1 / 2 for Humanoids.

  • Mechanism: Decompose the policy into asynchronous components: slow System 2 VLM for planning + fast System 1 for control, optionally with sparse mixture-of-experts in either.
  • Exemplars: GR00T N1/N1.5/N1.6, Hi-Robot, ฯ€0.5 / ฯ€0.7 (subtask + action expert), HiMoE-VLA, WholeBodyVLA, AdaMoE, RoboDual, OpenHelix, VITA-VLA (reverse distillation), Steerable Policies (Feb 2026 โ€” replaces NL S2/S1 interface with a 5-level multi-modal command vocabulary), DuoCore-FS (Dec 2025 โ€” Astribot's truly-parallel fast-slow with a written-and-read latent buffer + whole-body 3-stream RVQ-VAE tokenizer; 32.3 Hz on 25-DoF mobile dual-arm with 3B PaliGemma slow side), Galaxea + G0 (ICRA 2026 โ€” Qwen2.5-VL System-2 planner + PaliGemma-3B flow-matching System-1 actor; its evidence cuts against cross-embodiment scaling: a single-embodiment 500 h / 100 K-trajectory dataset drives few-shot generalization more than raw OXE scale, and the System-2 G0-VLM hits 83.3% on Table Bussing vs 32.0% for Gemini-2.5-pro). NeurIPS 2025 additions: Fast-in-Slow (S1 embedded inside S2 via shared parameters, not cascaded), ChatVLA-2 (Dynamic MoE that routes reasoning vs. action), ThinkAct (NVIDIA โ€” MLLM plans rewarded via RL, compressed to visual latent for action head), VLA-OS (controlled study that validates Hierarchical > Integrated > Action-Only).
  • Pros: โœ… Decouples control frequency from reasoning frequency ยท โœ… capacity scaling via sparse experts ยท โœ… best empirical results on long-horizon tasks.
  • Cons: โŒ Coordination between systems is finicky ยท โŒ sparsity training is unstable ยท โŒ MoE memory footprint ยท โŒ hand-designed inter-tier interface.
  • Pick when: Long-horizon / multi-embodiment / you can afford the coordination complexity.

Category G โ€” Reasoning-augmented (CoT)

  • Mechanism: Emit natural-language or symbolic intermediate tokens (plans, affordances, trajectory sketches, point-traces) before or alongside action tokens.
  • Exemplars: ECoT / ECoT-Lite, Hybrid Training, CoT-VLA, Embodied-R1, InstructVLA, MolmoAct, Vlaser, VLA-R1, CoA-VLA, ACoT-VLA, TraceVLA, dVLA (CoT channel), Steerable Policies (CoT reasoner emits grounded steering primitives โ€” points / traces โ€” alongside text rationales). NeurIPS 2025 additions: Chain-of-Action (ByteDance โ€” backward trajectory AR from goal keyframe), ThinkAct (RL-shaped MLLM plans as visual latent), Robot-R1 (R1-style RL on keypoint reasoning โ€” 7B beats GPT-4o on low-level spatial), DreamVLA (multi-modal forecasting as implicit CoT). ICRA 2026 addition: VLA-Reasoner moves the reasoning from emitted tokens to test-time search โ€” online MCTS over candidate action chunks, each branch rolled out through a 600M action-aware world model (iVideoGPT) and scored by an offline value head, so imagined futures are the "rationales." It is plug-and-play on a frozen policy and lifts OpenVLA real-world 22%โ†’41% and Octo-Small LIBERO 26.5%โ†’37.3% without retraining.
  • Pros: โœ… Interpretable ยท โœ… works with few demos ยท โœ… composes with any action-decoder category (A/B/C/D).
  • Cons: โŒ Longer sequences โ†’ slower inference (until dVLA solved this for D) ยท โŒ reasoning quality is weakly supervised ยท โŒ hallucinated reasoning decouples from actions.
  • Pick when: Long-horizon / need interpretability / few-shot adaptation.

Category H โ€” Tokenizer-centric efficiency

  • Mechanism: Optimize the action representation โ€” learned vector quantizers, DCT compression, B-spline bases, quantile binning โ€” to shrink tokens per action or enable structured parallel decoding. Orthogonal to categories A / D.
  • Exemplars: FAST (RSS 2025), FASTER (RVQ + DCT loss), OmniSAT (B-spline), HyperVLA (hypernetwork), AutoQVLA (channel-aware quantization).
  • Pros: โœ… Large latency / training-cost wins ยท โœ… enables entirely new decoding strategies ยท โœ… stacks with A or D.
  • Cons: โŒ Tokenizer quality is the new bottleneck ยท โŒ harder to interpret ยท โŒ can clash with cross-embodiment transfer.
  • Pick when: You're cost / latency constrained and already in category A or D.

Category I โ€” Small / efficient VLAs

  • Mechanism: Architectural choices targeting parameter / latency / single-GPU budgets โ€” smaller backbones, Mamba SSMs, layer pruning, distillation.
  • Exemplars: TinyVLA (2409.12514), SmolVLA (<0.5B), RoboMamba (SSM), FLOWER (950M flow-matching), NORA (Qwen-2.5-VL-3B), Lite VLA (CPU edge), ChatVLA, CogACT. NeurIPS 2025 additions: CogVLA (2508.21046 โ€” FiLM-based instruction-driven routing + token pruning; 97.4% LIBERO, 2.5ร— training / 2.8ร— inference speedup over OpenVLA). ICRA 2026 addition: LightVLA (H+I โ€” parameter-free, hyperparameter-free differentiable visual-token pruning via Gumbel-softmax over cross-attention saliency, ~78 tokens retained; โˆ’59.1% FLOPs / โˆ’38.2% latency while raising OpenVLA-OFT success 94.5%โ†’97.4%, where VLM-oriented pruners like FlashVLA/SP-VLA/VLA-Cache collapse to ~74%). It inverts the efficiency/accuracy trade-off by making the keep/drop decision performance-driven.
  • Pros: โœ… Deployable on edge / consumer HW ยท โœ… democratizes research ยท โœ… fast iteration.
  • Cons: โŒ Performance gap on hardest tasks ยท โŒ less cross-embodiment headroom.
  • Pick when: Edge deployment, consumer robot, or limited GPU budget.

Category J โ€” Sensory-augmented input (native 3D / tactile / event)

Contrast with Category L (3D-foundation-aligned). J = the robot has a non-RGB sensor (depth camera, GelSight tactile, event camera) and the VLA ingests it natively. L = the robot has only RGB but training aligns features to a 3D foundation model. Different deployment costs, different data-collection requirements.

  • Mechanism: Modify the input side, not the action decoder โ€” feed the policy point clouds, tactile, event-camera streams, or depth maps from real sensors.
  • Exemplars:
    • Native 3D point-cloud input: PointVLA (3D point-cloud features into frozen VLA), DexMove (ICLR-2026-DexMove โ€” R-Tac tactile sensor + 3D contact perception), Avi (input and output side โ€” also E5)
    • Tactile / force: Tactile-VLA, VLA-Touch (pretrained tactile-language), TaF-VLA (tactile + 6-axis force/torque), OmniVTLA (vision-tactile-language-action), DexMove (R-Tac 120 FPS markers + Farneback flow), FD-VLA (ICRA 2026 โ€” force awareness without a force sensor: a Force Distillation Module distills a force token from vision+state, aligned at training time against the latent of a real F/T signal, so the sensor is needed only during training; the distilled token reportedly outperforms the raw sensor reading โ€” recasting force sensing as a learnable-representation problem rather than a hardware requirement). โ†’ For the full touch-centric taxonomy (how tactile enters the policy, sensor hardware, ICRA/ICLR/CVPR 2026 trends, limitations), see the Tactile VLA cross-paper review.
    • Event cameras: E-VLA (2604.04834 / 2026) โ€” first event-camera VLA for dark/blurred scenes
    • Depth + RGB-D: AugVLA-3D
  • Pros: โœ… Unlocks contact-rich manipulation ยท โœ… direct geometric/force grounding ยท โœ… closer to real robot sensing stack ยท โœ… tactile is the only path to deformable / slip / friction-sensitive tasks.
  • Cons: โŒ Extra sensor pipelines + data collection ยท โŒ less internet-pretraining benefit for non-RGB modalities (no web-scale tactile data) ยท โŒ harder to scale cross-embodiment ยท โŒ sensor calibration / synchronization is a real deployment cost.
  • Pick when: Contact-rich / deformable / 3D-critical tasks where RGB+language alone fails. Combine with L if you have both sensors and foundation-model alignment.

Category K โ€” Hybrid AR + diffusion

  • Mechanism: Combine autoregressive next-token prediction with diffusion denoising in a single policy โ€” interleave the two objectives or ensemble their outputs.
  • Exemplars: HybridVLA (diffusion interleaved into AR; +14โ€“19% over SOTA), AR-VLA (AR over continuous action chunks), MMaDA-VLA, "Unifying Diffusion and Autoregression" (ICLR 2026 OpenReview H1KDMNOKQn).
  • Pros: โœ… Captures multimodal distribution (diffusion) without giving up token-level training (AR) ยท โœ… action-ensemble improves robustness.
  • Cons: โŒ Architectural complexity ยท โŒ two objectives to tune ยท โŒ unclear when to pick over A or D alone.
  • Pick when: You already have AR baseline and want diffusion's multimodality without migrating.

Category L โ€” 3D-foundation-aligned VLAs (new in this rebuild)

Why it's its own category, distinct from J. Category J (sensory-augmented) is about the input modality: the robot has a depth camera / tactile sensor / event camera, and the VLA ingests it. Category L is about the representation prior: the VLA only sees RGB+language at deploy, but during training its features are explicitly aligned to a 3D foundation model (VGGT, ESM, Sonata/PTv3, DINOv2-3D, SuperPoint). This is a 2025โ€“2026 thread that bypasses the data-collection cost of native-3D sensing while keeping the geometric grounding.

  • Mechanism: Train the VLA to predict / align with outputs of a frozen 3D foundation model โ€” depth maps, point clouds, part-aware feature fields, SE(3) transformations โ€” typically via an auxiliary loss at a specific VLM layer or a parallel head. Action prediction is otherwise standard (flow-matching, AR, or discrete diffusion).
  • Exemplars:
    • Spatial Forcing โ€” implicit alignment to VGGT (Wang et al. 2025) at VLM layer 24 of 32; LIBERO 98.5% average, 5.9ร— data efficiency vs the OpenVLA-OFT base
    • Spatially Guided Training (ST4VLA) โ€” Qwen2.5-VL-3B + DiT actor + 8.7 MB querying transformer with 0.5 gradient decay; SimplerEnv 66 โ†’ 84% on Visual Aggregation
    • Spatial-to-Actions / FALCON โ€” Kosmos-2 + 1.0B Embodied Spatial Module; stochastic Bernoulli p=0.66 conditioning of depth/pose; full CALVIN + SimplerEnv tables
    • EquAct โ€” SE(3)-equivariant multi-task transformer + iFiLM language conditioning; 18 RLBench tasks (Northeastern)
    • PA3FF โ€” Sonata/PTv3 part-aware dense 3D feature field + SigLIP semantic supervision; PartInstruct 28.79 vs GenDP 19.36, PartNetE 70.6 mAP50 zero-shot
    • SpatialVLA (RSS 2025) โ€” 3D Egocentric Position Encoding + Adaptive Spatial Grids derived from foundation 3D priors
    • GeoVLA (2508.09071 / 2025) โ€” parallel 2D VLM + Point Embedding Network derived from RGB
    • BridgeVLA (2506.07961 / NeurIPS 2025) โ€” projects 3D to multi-view 2D heatmaps for unified I/O; 96.8% real on 10 tasks with 3 trajectories each
    • Evaluation companion: Seeing Across Views (MV-RoboBench) shows even GPT-5 reaches only 56.4% vs human 91.0% on multi-view spatial reasoning โ€” the gap is the motivation for this whole category
  • Pros: โœ… Geometric grounding without 3D sensors at deploy ยท โœ… stacks with any action head ยท โœ… 5โ€“10ร— data efficiency reported across the cluster ยท โœ… benefits scale with 3D foundation model quality, not robot data.
  • Cons: โŒ Depends on a strong 3D foundation model (VGGT, Sonata, ESM, etc.) โ€” quality is the ceiling ยท โŒ alignment layer choice is empirical (Spatial Forcing settled on layer 24/32 by sweep) ยท โŒ adds compute during training even though inference is RGB-only.
  • Pick when: You want geometric grounding but cannot deploy depth sensors, or you want to multiply RGB sample efficiency.

Category M โ€” Latent-action / video-pretraining-derived VLAs (promoted from "Bonus" given 2025โ€“2026 maturity)

  • Mechanism: Extract latent action tokens from unlabeled video, then ground them to robot actions โ€” cleanly decouples "what" (latent) from "how" (per-embodiment decoder). Multiple sub-flavors: VQ-VAE quantization (LAPA), DINO-space task-centric (UniVLA), enhanced latent with proprio FDM (villa-X), Perceiver-resampler queries (BayesianVLA).
  • Exemplars:
    • LAPA (ICLR 2025) โ€” first unsupervised VLA pretraining via VQ-VAE latent actions from video
    • UniVLA latent (2505.06111) โ€” task-centric latent actions in DINO space, ~1/20 compute vs OpenVLA
    • villa-X โ€” PaliGemma 3B + 18-layer ACT-latent + 18-layer ACT-robot; proprio-FDM with 50%/50% stochastic masking; SIMPLER Google 77.7% / WidowX 62.5% vs ฯ€0 58.7%; 1.6M robot trajectories + 3.6M human-video clips
    • BayesianVLA (2601.15197 / 2026) โ€” Perceiver-resampler-style Latent Action Queries + PMI objective
    • XR-1 โ€” Unified Vision-Motion Codes (UVMC) as latent action representation for cross-embodiment
    • NavFoM (latent token application) โ€” Embodied Navigation Foundation Model uses TVI (Temporal-Viewpoint Indicator) tokens as a latent embodiment-context representation
    • GR00T's "LAPA embodiment" head โ€” pre-quantized VQ-VAE latents as pseudo-actions for unlabeled human ego-video (in the GR00T series data pyramid base layer)
  • Pros: โœ… Unsupervised pretraining from web video at scale ยท โœ… cross-embodiment transfer is natural (see Review-Cross-Embodiment ยง4.10) ยท โœ… strong sample efficiency (UniVLA ~1/20 compute vs. OpenVLA, villa-X transfers to XHand 12-DoF dexterous w/o dexterous pretraining data) ยท โœ… a clean separation of "intent" (latent) and "execution" (per-embodiment decoder) for the right tasks.
  • Cons: โŒ Latent quality is the ceiling โ€” and quality depends on the video corpus ยท โŒ fine-grained contact often doesn't fit a low-dim latent ยท โŒ task-centric latents collapse on busy but task-irrelevant video ยท โŒ a separate latent codebook complicates the training pipeline ยท โŒ LBM Co-training Study empirically reports that latent-action variants of human-video co-training do not help at LBM scale โ€” the gain comes from text captions, not from the latent codebook (Modality 4 verdict).
  • Pick when: You have abundant unlabeled video and limited robot data, and you can tolerate the latent-codebook engineering cost. Be aware of the LBM-Cotraining caveat before committing.

โ† Back to VLA Architectures review