Review VLA Architecture Papers - Heungwoo/research GitHub Wiki
VLA Architectures β The Papers (companion to Review-VLA-Architecture)
Split out of the main VLA Architectures review for faster GitHub-wiki rendering. This is the per-family paper list (Β§3 of that review).
Ordered roughly by architectural family. Bold = has a dedicated wiki page.
| Paper | Venue/year | Cat | Headline |
|---|---|---|---|
| RT-X / Open X-Embodiment | RSS 2024 | A | Unified discrete EE-pose tokens across 22 robots |
| OpenVLA | CoRL 2024 | A | 7B open AR VLA on OXE |
| Ο0-FAST | RSS 2025 | A+H | AR action tokens via FAST DCT tokenizer |
| VLA-0 | 2510.13054 / 2025 (NVIDIA) | A | Zero architectural change β actions as text; beats Ο0.5-KI, OpenVLA-OFT, GR00T-N1 on LIBERO |
| RT-2 | CoRL 2023 | A+G | Web-VLM β robot via co-fine-tune |
| Ο0 | RSS 2025 | B | First flow-matching action expert bolted onto a VLM |
| [Ο0.5](/Heungwoo/research/wiki/CoRL-2025-pi05) | CoRL 2025 Oral | B+F | Hierarchical subtask head + flow-matching expert |
| [Ο0.6](/Heungwoo/research/wiki/PI-pi06) | Nov 2025 | B | Gemma3-4B + Knowledge Insulation + optional metadata |
| [Ο0.7](/Heungwoo/research/wiki/PI-pi07) | Apr 2026 | B+F+G | MEM history + subgoal-image world model + metadata CFG |
| [RFS](/Heungwoo/research/wiki/ICLR-2026-RFS) | ICLR 2026 | B | Residual Flow Steering β residual flow head on frozen base |
| FLOWER | 2509.04996 / CoRL 2025 | B+I | 950M flow-matching VLA, 200 H100-hr pretrain |
| [Genesis AI GENE-26.5](/Heungwoo/research/wiki/Review-Genesis-GENE) | genesis.ai blog / May 2026 | B+J | Flow-matching unified VLA Β· 200,000+ h glove+ego+teleop+web Β· 3 ms / 500 Hz Β· closed weights Β· no arXiv |
| [Qwen-VLA](/Heungwoo/research/wiki/Review-Qwen-VLA) | 2605.30280 / May 2026 (Qwen Team / Alibaba) | B | Qwen3.5-4B + 1.15B DiT flow expert via concatenation + joint self-attention (a third Cat B wiring, distinct from Ο same-stack-MoE / GR00T cross-attn / LBM adaLN). Four-stage recipe (T2A β CPT β SFT β PPO RL with analytic flow-matching log-prob via ODEβSDE). Single generalist over 11 embodiments + human MANO + navigation: 97.9 LIBERO / 73.7 Simpler-WidowX / 86.1-87.2 RoboTwin / 76.9 OOD ALOHA / 26.6 zero-shot DOMINO (beats fine-tuned PUMA 17.2). No discrete action tokens, no FAST CE head β aligned with LBM. |
| RDT-1B | ICLR 2025 | C | 1.2B diffusion transformer, Physically Interpretable Unified Action Space |
| [DexVLA](/Heungwoo/research/wiki/CoRL-2025-DexVLA) | CoRL 2025 | C | ~1B plug-in diffusion action expert across arms/dex hands/bimanual |
| Diffusion Policy | RSS 2023 | C | The original continuous-diffusion manipulation policy |
| RoboDual | 2410.08001 / 2024β2025 | C+F | Diffusion-transformer specialist conditioned on VLA generalist; +26.7% real |
| [Discrete Diffusion VLA](/Heungwoo/research/wiki/ICLR-2026-Discrete-Diffusion-VLA) | ICLR 2026 | D | 96.3% LIBERO; single transformer, cross-entropy, adaptive unmasking |
| [Unified Diffusion VLA](/Heungwoo/research/wiki/ICLR-2026-Unified-Diffusion-VLA) | ICLR 2026 | D+E | Joint discrete denoising of future frames + actions |
| [dVLA](/Heungwoo/research/wiki/ICLR-2026-dVLA) | ICLR 2026 | D+G | Discrete diffusion with multimodal CoT (text + image + action in parallel) |
| [DIVA & Fast-dVLA](/Heungwoo/research/wiki/ICLR-2026-DIVA-Fast-dVLA) | ICLR 2026 | D+H | Latency optimizations for discrete-diffusion VLAs |
| MMaDA-VLA | 2603.25406 | D | Large diffusion-LM VLA with unified instruction + generation |
| Dream-VL / Dream-VLA | 2512.22615 | D | Whole backbone is a diffusion language model |
| [Cosmos Policy](/Heungwoo/research/wiki/ICLR-2026-Cosmos-Policy) | ICLR 2026 | E1 | NVIDIA Cosmos video foundation + control tokens (auxiliary world-model loss) |
| [DreamGen](/Heungwoo/research/wiki/CoRL-2025-DreamGen) | CoRL 2025 | E2 | Policy training inside video-world-model rollouts (world model as data factory) |
| [Ctrl-World](/Heungwoo/research/wiki/ICLR-2026-Ctrl-World) | ICLR 2026 | E3 | Pose-conditioned controllable world model as environment for policy evaluation |
| [WMPO](/Heungwoo/research/wiki/ICLR-2026-WMPO) | ICLR 2026 | E3 | World-model-based policy optimization (GRPO on OpenSora rollouts; 53β70% Mimicgen) |
| [Genie Envisioner](/Heungwoo/research/wiki/ICLR-2026-Genie-Envisioner) | ICLR 2026 | E1+E2 | LTX-Video 2B + GE-Act 160M; 5 Hz video / 30 Hz action async; 1M-episode AgiBot-World-Beta |
| [ViPRA](/Heungwoo/research/wiki/ICLR-2026-ViPRA) | ICLR 2026 | E2 | NSVQ 8-codebook + DINOv2 video pretrain β flow head; +16% SIMPLER, 22 Hz |
| [Vid2World](/Heungwoo/research/wiki/ICLR-2026-Vid2World) | ICLR 2026 | E1 | DynamiCrafter 1.1B + Diffusion Forcing + Causal Action Injection; CS:GO FVD β71%, RECON nav |
| [Geometry-aware 4D Video](/Heungwoo/research/wiki/ICLR-2026-Geometry-4D-Video) | ICLR 2026 | E1+J | Geometry-grounded video gen (FVD/AbsRel/Ξ΄β) + flow-matching policy |
| WorldVLA | 2506.21539 / 2025 | E1 | VLA + world model co-training |
| DreamVLA | 2507.04447 / 2025 | E1+G | Self-reflective; predicts dynamic regions / depth / semantics alongside action |
| mimic-video (VAM) | 2512.15692 / 2025 | E4 | Replaces VLM with pretrained video-generation model; order-of-magnitude better sample efficiency |
| DiT4DiT | 2603.10448 / 2026 | E4 | Dual-DiT video-action model; no VLM autoregressive backbone |
| S-VAM | 2603.16195 / 2026 | E4 | Shortcut VAM via self-distilling foresight |
| Avi | 2510.21746 / NeurIPS 2025 Workshop (Embodied World Models) | E5+J | Predicts future point cloud, extracts action via IK β skips action tokens entirely (geometry-first) |
| GR00T N1 / N1.5 / N1.6 | 2503.14734 / 2025 | F | Dual-system: Eagle-2 VLM + DiT; open humanoid foundation |
| Hi-Robot | 2025 | F | Hierarchical planner + low-level controller |
| [HiMoE-VLA](/Heungwoo/research/wiki/ICLR-2026-HiMoE-VLA) | ICLR 2026 | F | Hierarchical MoE routing |
| [WholeBodyVLA](/Heungwoo/research/wiki/ICLR-2026-WholeBodyVLA) | ICLR 2026 | F | Unified latent β coordinated humanoid base/arms/hands |
| AdaMoE | 2510.14300 / ICLR 2026 | F | Action-specialized MoE; sparsifies dense action expert FFN |
| OpenHelix | 2505.03912 / 2025 | F | Open-source dual-system reference model |
| VITA-VLA | 2510.09607 / 2025 | F | Distills small action model into a 7B VLM (reverse direction) |
| [DuoCore-FS](/Heungwoo/research/wiki/Review-DuoCore-FS) | 2512.20188 / Dec 2025 (Astribot) | F+H | Truly parallel fast-slow whole-body VLA. Ο0-FAST on PaliGemma-3B @ 1β3 Hz + Pi0-small flow-matching @ 25β30 Hz, decoupled by a bridge buffer of instruction + learnable fusion-query embeddings (fusion-query params trained through fast-side loss = differentiable interface). Whole-body action tokenizer = 3-stream residual VQ-VAE (pos / 6D-SO(3) / gripper) Γ codebook 1024 with geodesic SO(3) loss β 36 fixed tokens vs FAST's avg 81 / max 205. Trained jointly E2E with Ξ βΌ U[0,25]-frame delay sampling. 32.3 Hz vs Ο0 12.5 Hz on RTX 4090; 90% vs 85% in-distribution; 50% vs 10% OOD; 42.9% vs 14.3% language following on a popcorn-kiosk task (1,780 traj / 10.22 h, Astribot S1 25-DoF mobile dual-arm). Closed implementation (commercial via Astribot). Differentiates explicitly from FiS-VLA / OpenHelix (fixed-ratio scheduling), Helix (closed), Hume (cascaded, not E2E). |
| π [Galaxea + G0](/Heungwoo/research/wiki/ICRA-2026-Galaxea-G0) | 2509.00576 / ICRA 2026 | F | Dual-system: Qwen2.5-VL planner (System-2) + PaliGemma-3B flow-matching actor (System-1), trained on a 500 h / 100 K-trajectory open-world dataset |
| ECoT / ECoT-Lite | Berkeley 2024β2025 | G | Embodied chain-of-thought |
| [Hybrid Training](/Heungwoo/research/wiki/ICLR-2026-Hybrid-Training) | ICLR 2026 | G | Makes ECoT skippable at deployment |
| MolmoAct | 2508.07917 / AllenAI 2025 | G | "Action Reasoning Models"; reason in 3D before acting |
| Vlaser | 2510.11027 / ICLR 2026 | G | Vlaser-6M dataset; synergistic embodied reasoning |
| CoT-VLA | CVPR 2025 | G+K | Visual CoT for VLAs (arguably hybrid AR+diffusion) |
| CoA-VLA | ICCV 2025 | G | Chain-of-Affordance |
| ACoT-VLA | 2601.11404 / CVPR 2026 | G | Action-CoT with Explicit + Implicit Reasoners |
| VLA-R1 | 2510.01623 / 2025 | G | R1-style RL over reasoning + execution |
| π [VLA-Reasoner](/Heungwoo/research/wiki/ICRA-2026-VLA-Reasoner) | 2509.22643 / ICRA 2026 | G | Plug-and-play online MCTS with a world model + value net at test time; OpenVLA real-world 22% β 41% |
| [Steerable Policies](/Heungwoo/research/wiki/Review-Steerable-Policies) | 2602.13193 / Feb 2026 | F+G | 5-level steering vocabulary (subtask + atomic motion + point + trace + hybrid) replaces NL-only S2/S1 interface; backbone-agnostic on OpenVLA + Ο0.5; off-the-shelf Gemini ICL drives the policy without high-level fine-tune |
| TraceVLA (ICLR-2025) | 2412.10345 | G | Visual trace prompting |
| [Embodied-R1](/Heungwoo/research/wiki/ICLR-2026-Embodied-R1) | ICLR 2026 | G | R1 RL on pointing primitives |
| [InstructVLA](/Heungwoo/research/wiki/ICLR-2026-InstructVLA) | ICLR 2026 | G | VLA instruction tuning; VLA-IT 650K |
| [FASTER](/Heungwoo/research/wiki/ICLR-2026-FASTER) | ICLR 2026 | H | RVQ + DCT loss action tokens |
| [OmniSAT](/Heungwoo/research/wiki/ICLR-2026-OmniSAT) | ICLR 2026 | H | B-spline action tokens |
| [HyperVLA](/Heungwoo/research/wiki/ICLR-2026-HyperVLA) | ICLR 2026 | H | Hypernetwork action decoder |
| [AutoQVLA](/Heungwoo/research/wiki/ICLR-2026-AutoQVLA) | ICLR 2026 | H | Channel-aware quantization |
| TinyVLA | 2409.12514 | I | Small data-efficient VLA |
| SmolVLA | 2506.01844 / HF-LeRobot 2025 | I | <0.5B params, single-GPU training |
| RoboMamba | 2406.04339 | I | Mamba SSM backbone; 0.1% policy-head params |
| π [LightVLA (token pruning)](/Heungwoo/research/wiki/ICRA-2026-Token-Pruning) | 2509.12594 / ICRA 2026 | H+I | Differentiable, parameter-free visual-token pruning (Gumbel-softmax over cross-attention saliency); β59% FLOPs / β38% latency and +2.9% success on OpenVLA-OFT |
| NORA | 2504.19854 / 2025 | I | Small generalist VLA on Qwen-2.5-VL-3B |
| Lite VLA | 2511.05642 / 2025 | I | CPU-bound edge deployment |
| ChatVLA | 2502.14420 | I | Unified multimodal understanding + control |
| CogACT | 2411.19650 | I+F | Cognition + action synergy |
| SpatialVLA | RSS 2025 | L | 3D Egocentric Position Encoding + Adaptive Spatial Grids (foundation 3D priors from RGB) |
| PointVLA | 2503.07511 / 2025 | J | 3D point-cloud features injected into frozen VLA |
| GeoVLA | 2508.09071 / 2025 | L | Parallel 2D VLM + Point Embedding Network (geometry features from RGB) |
| [Spatial Forcing](/Heungwoo/research/wiki/ICLR-2026-Spatial-Forcing) | ICLR 2026 | L | Implicit alignment to VGGT 3D foundation model at VLM layer 24; LIBERO 98.5% |
| [Spatially Guided Training (ST4VLA)](/Heungwoo/research/wiki/ICLR-2026-Spatially-Guided) | ICLR 2026 | L | Two-stage spatial pretraining + DiT actor; SimplerEnv 66β84% |
| [Spatial-to-Actions (FALCON)](/Heungwoo/research/wiki/ICLR-2026-Spatial-to-Actions) | ICLR 2026 | L | Kosmos-2 + ESM spatial foundation priors injected at action head |
| [EquAct](/Heungwoo/research/wiki/ICLR-2026-EquAct) | ICLR 2026 | L | SE(3)-equivariant transformer + iFiLM; 18 RLBench tasks |
| [PA3FF](/Heungwoo/research/wiki/ICLR-2026-PA3FF) | ICLR 2026 | L | Sonata/PTv3 part-aware 3D feature field + SigLIP semantic supervision |
| BridgeVLA | 2506.07961 / NeurIPS 2025 | L | Projects 3D to multi-view 2D heatmaps for unified I/O; 96.8% real on 10 tasks |
| [Seeing Across Views](/Heungwoo/research/wiki/ICLR-2026-Seeing-Across-Views) | ICLR 2026 | L (eval) | MV-RoboBench: VLMs struggle with multi-view spatial reasoning (best ~56% vs human 91%) |
| Tactile-VLA | 2507.09160 | J | Tactile + Ο0-style backbone + hybrid position-force |
| VLA-Touch | 2507.17294 | J | Pretrained tactile-language + diffusion controller |
| TaF-VLA | 2601.20321 / 2026 | J | Tactile + 6-axis force/torque |
| OmniVTLA | 2508.08706 / 2025 | J | Vision-Tactile-Language-Action |
| π [OmniVTA](/Heungwoo/research/wiki/Review-OmniVTA) | 2603.19201 / 2026 | J + WM | Visuo-tactile world model (predict contact evolution) + 60 Hz reflexive control; not a language-conditioned VLA |
| E-VLA | 2604.04834 / 2026 | J | Event-camera VLA for dark / blurred scenes |
| HybridVLA | 2503.10631 / Mar 2025 | K | Single LLM with diffusion denoising interleaved into AR |
| AR-VLA | 2603.10126 / 2026 | K | True AR over continuous action chunks |
| LAPA | 2410.11758 / ICLR 2025 | (latent) | First unsupervised VLA pretraining via latent actions from video |
| UniVLA (latent actions) | 2505.06111 | (latent) | Task-centric latent actions in DINO space |
| villa-X | 2507.23682 | (latent) | Enhanced latent-action modeling |
| BayesianVLA | 2601.15197 / 2026 | (latent) | Perceiver-resampler-style Latent Action Queries + PMI objective |
| VLM4VLA | 2601.03309 / ICLR 2026 | (analysis) | See Review-VLM4VLA β VLM-bottleneck analysis |
| [Knowledge Insulation](/Heungwoo/research/wiki/NeurIPS-2025-Knowledge-Insulation) | NeurIPS 2025 Spotlight, PI | B (training recipe) | Gradient-insulated VLM + continuous action expert; the published Ο0.5-KI β Ο0.6 β Ο0.7 bridge |
| [ChatVLA-2](/Heungwoo/research/wiki/NeurIPS-2025-ChatVLA-2) | NeurIPS 2025 | F | Dynamic MoE that preserves VLM pretraining during action fine-tuning; 82.7% on open-world math-matching |
| [DreamVLA](/Heungwoo/research/wiki/NeurIPS-2025-DreamVLA) | NeurIPS 2025 | E+G | Multi-modal world-knowledge forecasting (dynamic regions + depth + geom + seg) as inverse-dynamics signal |
| [VLA-OS](/Heungwoo/research/wiki/NeurIPS-2025-VLA-OS) | NeurIPS 2025 | F (meta-study) | Controlled architecture ablation: Hierarchical > Integrated > Action-Only; visual-grounded > language-grounded planning |
| [Fast-in-Slow](/Heungwoo/research/wiki/NeurIPS-2025-Fast-in-Slow) | NeurIPS 2025 | F | Embedded (not cascaded) System-1-in-System-2 via shared parameters; 117.7 Hz control |
| [ThinkAct](/Heungwoo/research/wiki/NeurIPS-2025-ThinkAct) | NeurIPS 2025 / NVIDIA | F+G | MLLM plans rewarded by RL (goal completion + trajectory consistency); plan β visual latent β action |
| [Chain-of-Action](/Heungwoo/research/wiki/NeurIPS-2025-Chain-of-Action) | NeurIPS 2025 / ByteDance | A+G | Backward trajectory AR from goal keyframe β global-to-local action CoT |
| [Real-Time Chunking](/Heungwoo/research/wiki/NeurIPS-2025-Real-Time-Chunking) | NeurIPS 2025 / PI + Berkeley | B+H | Async chunk inpainting β no pause at boundaries; works on any diffusion/flow VLA |
| [ReinFlow](/Heungwoo/research/wiki/NeurIPS-2025-ReinFlow) | NeurIPS 2025 | B (RL) | Learnable noise injection gives exact likelihoods β online RL on Ο0/Ο0.5/GR00T-N1.5 |
| [VideoVLA](/Heungwoo/research/wiki/NeurIPS-2025-VideoVLA) | NeurIPS 2025 | E (VAM) | Multimodal DiT jointly predicts action chunks + future video; imagined futures correlate with success |
| CogVLA | 2508.21046 / NeurIPS 2025 | I+H | FiLM-based instruction-driven routing + token pruning; 2.5Γ training / 2.8Γ inference speedup |
| BridgeVLA | 2506.07961 / NeurIPS 2025 | J (3D) | Projects 3D to multi-view 2D; I/O unified as 2D heatmaps; 96.8% real on 10 tasks w/ 3 trajectories each |
| DynaGuide | 2506.13922 / NeurIPS 2025 / Stanford | C | Latent-dynamics-guided inference-time steering of pretrained diffusion policies; up-weights rare behaviors |
β Back to VLA Architectures review