Review VLA Architecture Papers - Heungwoo/research GitHub Wiki

VLA Architectures β€” The Papers (companion to Review-VLA-Architecture)

Split out of the main VLA Architectures review for faster GitHub-wiki rendering. This is the per-family paper list (Β§3 of that review).

Ordered roughly by architectural family. Bold = has a dedicated wiki page.

Paper Venue/year Cat Headline
RT-X / Open X-Embodiment RSS 2024 A Unified discrete EE-pose tokens across 22 robots
OpenVLA CoRL 2024 A 7B open AR VLA on OXE
Ο€0-FAST RSS 2025 A+H AR action tokens via FAST DCT tokenizer
VLA-0 2510.13054 / 2025 (NVIDIA) A Zero architectural change β€” actions as text; beats Ο€0.5-KI, OpenVLA-OFT, GR00T-N1 on LIBERO
RT-2 CoRL 2023 A+G Web-VLM β†’ robot via co-fine-tune
Ο€0 RSS 2025 B First flow-matching action expert bolted onto a VLM
[Ο€0.5](/Heungwoo/research/wiki/CoRL-2025-pi05) CoRL 2025 Oral B+F Hierarchical subtask head + flow-matching expert
[Ο€0.6](/Heungwoo/research/wiki/PI-pi06) Nov 2025 B Gemma3-4B + Knowledge Insulation + optional metadata
[Ο€0.7](/Heungwoo/research/wiki/PI-pi07) Apr 2026 B+F+G MEM history + subgoal-image world model + metadata CFG
[RFS](/Heungwoo/research/wiki/ICLR-2026-RFS) ICLR 2026 B Residual Flow Steering β€” residual flow head on frozen base
FLOWER 2509.04996 / CoRL 2025 B+I 950M flow-matching VLA, 200 H100-hr pretrain
[Genesis AI GENE-26.5](/Heungwoo/research/wiki/Review-Genesis-GENE) genesis.ai blog / May 2026 B+J Flow-matching unified VLA Β· 200,000+ h glove+ego+teleop+web Β· 3 ms / 500 Hz Β· closed weights Β· no arXiv
[Qwen-VLA](/Heungwoo/research/wiki/Review-Qwen-VLA) 2605.30280 / May 2026 (Qwen Team / Alibaba) B Qwen3.5-4B + 1.15B DiT flow expert via concatenation + joint self-attention (a third Cat B wiring, distinct from π same-stack-MoE / GR00T cross-attn / LBM adaLN). Four-stage recipe (T2A → CPT → SFT → PPO RL with analytic flow-matching log-prob via ODE→SDE). Single generalist over 11 embodiments + human MANO + navigation: 97.9 LIBERO / 73.7 Simpler-WidowX / 86.1-87.2 RoboTwin / 76.9 OOD ALOHA / 26.6 zero-shot DOMINO (beats fine-tuned PUMA 17.2). No discrete action tokens, no FAST CE head — aligned with LBM.
RDT-1B ICLR 2025 C 1.2B diffusion transformer, Physically Interpretable Unified Action Space
[DexVLA](/Heungwoo/research/wiki/CoRL-2025-DexVLA) CoRL 2025 C ~1B plug-in diffusion action expert across arms/dex hands/bimanual
Diffusion Policy RSS 2023 C The original continuous-diffusion manipulation policy
RoboDual 2410.08001 / 2024–2025 C+F Diffusion-transformer specialist conditioned on VLA generalist; +26.7% real
[Discrete Diffusion VLA](/Heungwoo/research/wiki/ICLR-2026-Discrete-Diffusion-VLA) ICLR 2026 D 96.3% LIBERO; single transformer, cross-entropy, adaptive unmasking
[Unified Diffusion VLA](/Heungwoo/research/wiki/ICLR-2026-Unified-Diffusion-VLA) ICLR 2026 D+E Joint discrete denoising of future frames + actions
[dVLA](/Heungwoo/research/wiki/ICLR-2026-dVLA) ICLR 2026 D+G Discrete diffusion with multimodal CoT (text + image + action in parallel)
[DIVA & Fast-dVLA](/Heungwoo/research/wiki/ICLR-2026-DIVA-Fast-dVLA) ICLR 2026 D+H Latency optimizations for discrete-diffusion VLAs
MMaDA-VLA 2603.25406 D Large diffusion-LM VLA with unified instruction + generation
Dream-VL / Dream-VLA 2512.22615 D Whole backbone is a diffusion language model
[Cosmos Policy](/Heungwoo/research/wiki/ICLR-2026-Cosmos-Policy) ICLR 2026 E1 NVIDIA Cosmos video foundation + control tokens (auxiliary world-model loss)
[DreamGen](/Heungwoo/research/wiki/CoRL-2025-DreamGen) CoRL 2025 E2 Policy training inside video-world-model rollouts (world model as data factory)
[Ctrl-World](/Heungwoo/research/wiki/ICLR-2026-Ctrl-World) ICLR 2026 E3 Pose-conditioned controllable world model as environment for policy evaluation
[WMPO](/Heungwoo/research/wiki/ICLR-2026-WMPO) ICLR 2026 E3 World-model-based policy optimization (GRPO on OpenSora rollouts; 53β†’70% Mimicgen)
[Genie Envisioner](/Heungwoo/research/wiki/ICLR-2026-Genie-Envisioner) ICLR 2026 E1+E2 LTX-Video 2B + GE-Act 160M; 5 Hz video / 30 Hz action async; 1M-episode AgiBot-World-Beta
[ViPRA](/Heungwoo/research/wiki/ICLR-2026-ViPRA) ICLR 2026 E2 NSVQ 8-codebook + DINOv2 video pretrain β†’ flow head; +16% SIMPLER, 22 Hz
[Vid2World](/Heungwoo/research/wiki/ICLR-2026-Vid2World) ICLR 2026 E1 DynamiCrafter 1.1B + Diffusion Forcing + Causal Action Injection; CS:GO FVD βˆ’71%, RECON nav
[Geometry-aware 4D Video](/Heungwoo/research/wiki/ICLR-2026-Geometry-4D-Video) ICLR 2026 E1+J Geometry-grounded video gen (FVD/AbsRel/δ₁) + flow-matching policy
WorldVLA 2506.21539 / 2025 E1 VLA + world model co-training
DreamVLA 2507.04447 / 2025 E1+G Self-reflective; predicts dynamic regions / depth / semantics alongside action
mimic-video (VAM) 2512.15692 / 2025 E4 Replaces VLM with pretrained video-generation model; order-of-magnitude better sample efficiency
DiT4DiT 2603.10448 / 2026 E4 Dual-DiT video-action model; no VLM autoregressive backbone
S-VAM 2603.16195 / 2026 E4 Shortcut VAM via self-distilling foresight
Avi 2510.21746 / NeurIPS 2025 Workshop (Embodied World Models) E5+J Predicts future point cloud, extracts action via IK β€” skips action tokens entirely (geometry-first)
GR00T N1 / N1.5 / N1.6 2503.14734 / 2025 F Dual-system: Eagle-2 VLM + DiT; open humanoid foundation
Hi-Robot 2025 F Hierarchical planner + low-level controller
[HiMoE-VLA](/Heungwoo/research/wiki/ICLR-2026-HiMoE-VLA) ICLR 2026 F Hierarchical MoE routing
[WholeBodyVLA](/Heungwoo/research/wiki/ICLR-2026-WholeBodyVLA) ICLR 2026 F Unified latent β†’ coordinated humanoid base/arms/hands
AdaMoE 2510.14300 / ICLR 2026 F Action-specialized MoE; sparsifies dense action expert FFN
OpenHelix 2505.03912 / 2025 F Open-source dual-system reference model
VITA-VLA 2510.09607 / 2025 F Distills small action model into a 7B VLM (reverse direction)
[DuoCore-FS](/Heungwoo/research/wiki/Review-DuoCore-FS) 2512.20188 / Dec 2025 (Astribot) F+H Truly parallel fast-slow whole-body VLA. Ο€0-FAST on PaliGemma-3B @ 1–3 Hz + Pi0-small flow-matching @ 25–30 Hz, decoupled by a bridge buffer of instruction + learnable fusion-query embeddings (fusion-query params trained through fast-side loss = differentiable interface). Whole-body action tokenizer = 3-stream residual VQ-VAE (pos / 6D-SO(3) / gripper) Γ— codebook 1024 with geodesic SO(3) loss β€” 36 fixed tokens vs FAST's avg 81 / max 205. Trained jointly E2E with Ξ” ∼ U[0,25]-frame delay sampling. 32.3 Hz vs Ο€0 12.5 Hz on RTX 4090; 90% vs 85% in-distribution; 50% vs 10% OOD; 42.9% vs 14.3% language following on a popcorn-kiosk task (1,780 traj / 10.22 h, Astribot S1 25-DoF mobile dual-arm). Closed implementation (commercial via Astribot). Differentiates explicitly from FiS-VLA / OpenHelix (fixed-ratio scheduling), Helix (closed), Hume (cascaded, not E2E).
πŸ†• [Galaxea + G0](/Heungwoo/research/wiki/ICRA-2026-Galaxea-G0) 2509.00576 / ICRA 2026 F Dual-system: Qwen2.5-VL planner (System-2) + PaliGemma-3B flow-matching actor (System-1), trained on a 500 h / 100 K-trajectory open-world dataset
ECoT / ECoT-Lite Berkeley 2024–2025 G Embodied chain-of-thought
[Hybrid Training](/Heungwoo/research/wiki/ICLR-2026-Hybrid-Training) ICLR 2026 G Makes ECoT skippable at deployment
MolmoAct 2508.07917 / AllenAI 2025 G "Action Reasoning Models"; reason in 3D before acting
Vlaser 2510.11027 / ICLR 2026 G Vlaser-6M dataset; synergistic embodied reasoning
CoT-VLA CVPR 2025 G+K Visual CoT for VLAs (arguably hybrid AR+diffusion)
CoA-VLA ICCV 2025 G Chain-of-Affordance
ACoT-VLA 2601.11404 / CVPR 2026 G Action-CoT with Explicit + Implicit Reasoners
VLA-R1 2510.01623 / 2025 G R1-style RL over reasoning + execution
πŸ†• [VLA-Reasoner](/Heungwoo/research/wiki/ICRA-2026-VLA-Reasoner) 2509.22643 / ICRA 2026 G Plug-and-play online MCTS with a world model + value net at test time; OpenVLA real-world 22% β†’ 41%
[Steerable Policies](/Heungwoo/research/wiki/Review-Steerable-Policies) 2602.13193 / Feb 2026 F+G 5-level steering vocabulary (subtask + atomic motion + point + trace + hybrid) replaces NL-only S2/S1 interface; backbone-agnostic on OpenVLA + Ο€0.5; off-the-shelf Gemini ICL drives the policy without high-level fine-tune
TraceVLA (ICLR-2025) 2412.10345 G Visual trace prompting
[Embodied-R1](/Heungwoo/research/wiki/ICLR-2026-Embodied-R1) ICLR 2026 G R1 RL on pointing primitives
[InstructVLA](/Heungwoo/research/wiki/ICLR-2026-InstructVLA) ICLR 2026 G VLA instruction tuning; VLA-IT 650K
[FASTER](/Heungwoo/research/wiki/ICLR-2026-FASTER) ICLR 2026 H RVQ + DCT loss action tokens
[OmniSAT](/Heungwoo/research/wiki/ICLR-2026-OmniSAT) ICLR 2026 H B-spline action tokens
[HyperVLA](/Heungwoo/research/wiki/ICLR-2026-HyperVLA) ICLR 2026 H Hypernetwork action decoder
[AutoQVLA](/Heungwoo/research/wiki/ICLR-2026-AutoQVLA) ICLR 2026 H Channel-aware quantization
TinyVLA 2409.12514 I Small data-efficient VLA
SmolVLA 2506.01844 / HF-LeRobot 2025 I <0.5B params, single-GPU training
RoboMamba 2406.04339 I Mamba SSM backbone; 0.1% policy-head params
πŸ†• [LightVLA (token pruning)](/Heungwoo/research/wiki/ICRA-2026-Token-Pruning) 2509.12594 / ICRA 2026 H+I Differentiable, parameter-free visual-token pruning (Gumbel-softmax over cross-attention saliency); βˆ’59% FLOPs / βˆ’38% latency and +2.9% success on OpenVLA-OFT
NORA 2504.19854 / 2025 I Small generalist VLA on Qwen-2.5-VL-3B
Lite VLA 2511.05642 / 2025 I CPU-bound edge deployment
ChatVLA 2502.14420 I Unified multimodal understanding + control
CogACT 2411.19650 I+F Cognition + action synergy
SpatialVLA RSS 2025 L 3D Egocentric Position Encoding + Adaptive Spatial Grids (foundation 3D priors from RGB)
PointVLA 2503.07511 / 2025 J 3D point-cloud features injected into frozen VLA
GeoVLA 2508.09071 / 2025 L Parallel 2D VLM + Point Embedding Network (geometry features from RGB)
[Spatial Forcing](/Heungwoo/research/wiki/ICLR-2026-Spatial-Forcing) ICLR 2026 L Implicit alignment to VGGT 3D foundation model at VLM layer 24; LIBERO 98.5%
[Spatially Guided Training (ST4VLA)](/Heungwoo/research/wiki/ICLR-2026-Spatially-Guided) ICLR 2026 L Two-stage spatial pretraining + DiT actor; SimplerEnv 66β†’84%
[Spatial-to-Actions (FALCON)](/Heungwoo/research/wiki/ICLR-2026-Spatial-to-Actions) ICLR 2026 L Kosmos-2 + ESM spatial foundation priors injected at action head
[EquAct](/Heungwoo/research/wiki/ICLR-2026-EquAct) ICLR 2026 L SE(3)-equivariant transformer + iFiLM; 18 RLBench tasks
[PA3FF](/Heungwoo/research/wiki/ICLR-2026-PA3FF) ICLR 2026 L Sonata/PTv3 part-aware 3D feature field + SigLIP semantic supervision
BridgeVLA 2506.07961 / NeurIPS 2025 L Projects 3D to multi-view 2D heatmaps for unified I/O; 96.8% real on 10 tasks
[Seeing Across Views](/Heungwoo/research/wiki/ICLR-2026-Seeing-Across-Views) ICLR 2026 L (eval) MV-RoboBench: VLMs struggle with multi-view spatial reasoning (best ~56% vs human 91%)
Tactile-VLA 2507.09160 J Tactile + Ο€0-style backbone + hybrid position-force
VLA-Touch 2507.17294 J Pretrained tactile-language + diffusion controller
TaF-VLA 2601.20321 / 2026 J Tactile + 6-axis force/torque
OmniVTLA 2508.08706 / 2025 J Vision-Tactile-Language-Action
πŸ†• [OmniVTA](/Heungwoo/research/wiki/Review-OmniVTA) 2603.19201 / 2026 J + WM Visuo-tactile world model (predict contact evolution) + 60 Hz reflexive control; not a language-conditioned VLA
E-VLA 2604.04834 / 2026 J Event-camera VLA for dark / blurred scenes
HybridVLA 2503.10631 / Mar 2025 K Single LLM with diffusion denoising interleaved into AR
AR-VLA 2603.10126 / 2026 K True AR over continuous action chunks
LAPA 2410.11758 / ICLR 2025 (latent) First unsupervised VLA pretraining via latent actions from video
UniVLA (latent actions) 2505.06111 (latent) Task-centric latent actions in DINO space
villa-X 2507.23682 (latent) Enhanced latent-action modeling
BayesianVLA 2601.15197 / 2026 (latent) Perceiver-resampler-style Latent Action Queries + PMI objective
VLM4VLA 2601.03309 / ICLR 2026 (analysis) See Review-VLM4VLA β€” VLM-bottleneck analysis
[Knowledge Insulation](/Heungwoo/research/wiki/NeurIPS-2025-Knowledge-Insulation) NeurIPS 2025 Spotlight, PI B (training recipe) Gradient-insulated VLM + continuous action expert; the published Ο€0.5-KI β†’ Ο€0.6 β†’ Ο€0.7 bridge
[ChatVLA-2](/Heungwoo/research/wiki/NeurIPS-2025-ChatVLA-2) NeurIPS 2025 F Dynamic MoE that preserves VLM pretraining during action fine-tuning; 82.7% on open-world math-matching
[DreamVLA](/Heungwoo/research/wiki/NeurIPS-2025-DreamVLA) NeurIPS 2025 E+G Multi-modal world-knowledge forecasting (dynamic regions + depth + geom + seg) as inverse-dynamics signal
[VLA-OS](/Heungwoo/research/wiki/NeurIPS-2025-VLA-OS) NeurIPS 2025 F (meta-study) Controlled architecture ablation: Hierarchical > Integrated > Action-Only; visual-grounded > language-grounded planning
[Fast-in-Slow](/Heungwoo/research/wiki/NeurIPS-2025-Fast-in-Slow) NeurIPS 2025 F Embedded (not cascaded) System-1-in-System-2 via shared parameters; 117.7 Hz control
[ThinkAct](/Heungwoo/research/wiki/NeurIPS-2025-ThinkAct) NeurIPS 2025 / NVIDIA F+G MLLM plans rewarded by RL (goal completion + trajectory consistency); plan β†’ visual latent β†’ action
[Chain-of-Action](/Heungwoo/research/wiki/NeurIPS-2025-Chain-of-Action) NeurIPS 2025 / ByteDance A+G Backward trajectory AR from goal keyframe β€” global-to-local action CoT
[Real-Time Chunking](/Heungwoo/research/wiki/NeurIPS-2025-Real-Time-Chunking) NeurIPS 2025 / PI + Berkeley B+H Async chunk inpainting β€” no pause at boundaries; works on any diffusion/flow VLA
[ReinFlow](/Heungwoo/research/wiki/NeurIPS-2025-ReinFlow) NeurIPS 2025 B (RL) Learnable noise injection gives exact likelihoods β†’ online RL on Ο€0/Ο€0.5/GR00T-N1.5
[VideoVLA](/Heungwoo/research/wiki/NeurIPS-2025-VideoVLA) NeurIPS 2025 E (VAM) Multimodal DiT jointly predicts action chunks + future video; imagined futures correlate with success
CogVLA 2508.21046 / NeurIPS 2025 I+H FiLM-based instruction-driven routing + token pruning; 2.5Γ— training / 2.8Γ— inference speedup
BridgeVLA 2506.07961 / NeurIPS 2025 J (3D) Projects 3D to multi-view 2D; I/O unified as 2D heatmaps; 96.8% real on 10 tasks w/ 3 trajectories each
DynaGuide 2506.13922 / NeurIPS 2025 / Stanford C Latent-dynamics-guided inference-time steering of pretrained diffusion policies; up-weights rare behaviors

← Back to VLA Architectures review