π€ VLA / Robot AI Research Review Wiki
A living research wiki for Vision-Language-Action models Β· robot manipulation Β· humanoid intelligence Β· tactile & force interaction Β· world models .
π 396 pages
π 8 venues
π 16+ cross-paper reviews
π updated 2026-07-25
New here? Start with one of these:
π§ VLA Architectures β choose the model family & action decoder Β·
π Ο Series Evolution β the production-VLA baseline lineage Β·
π IROS 2026 Survey π Β· RSS 2026 Survey β the latest venue surveys Β·
π Latest Papers β preprint tracker (pre-publication reviews) Β·
πΈοΈ Knowledge Graph β the visual map
πΊ 1. Research Decision Map
Pick your question β open the fold-out for a 4-line orientation β follow the π deep-dive link for the full survey (trend arc Β· approach taxonomy Β· limitations). Every deep-dive page carries a matching "State of the Field" section, updated Aug 2026.
Which VLA architecture should we use? β flow matching dominates; ordered tokens revived AR
π Trend
AR tokens (2024) β flow-matching experts standard (2025) β flow universal in flagships + AR revived via OAT 's anytime prefix decoding (2026)
βοΈ Approaches
AR discrete tokens (LLM-native / lossy, slow) Β· flow-matching expert (precise, production-proven / no native log-probs) Β· discrete diffusion in-VLM (parallel decode / little traction)
β οΈ Open
No matched-scale comparison across the three families; latency rarely published
How should the VLM connect to the action expert? β evidence tilts to cross-attention, margins small
π Trend
2025 four-way standoff β first direct ablations in 2026: cross-attention wins in RobotManip (87.5 vs 87.0) and MM-DiT beats naive DiT in Ξ¨β
βοΈ Approaches
same-stack MoE+prefix-KV (tight / hard to retrofit) Β· cross-attention (decoupled, small experts / one-way) Β· concatenation (simplest / re-attends per step) Β· adaLN token (cheapest / bottleneck)
β οΈ Open
~0.5 pp margins on one benchmark; the two Qwen flagships disagree internally
How should attention / KV-cache be designed? β the runtime now drives the choice
π Trend
Design follows the runtime since 2025 (RTC/streaming/prefix-KV); 2026 adds linear-attention backbones (Qwen3.5) with zero control-quality ablation
βοΈ Approaches
prefix-KV + expert branch (cache reuse) Β· full re-attention per chunk (freshness) Β· streaming-token designs
β οΈ Open
Latency is the field's least-reported number; multi-view KV behavior is folklore
Should vision be built outside the VLM? β the ViT is the proven bottleneck; fix = action supervision
π Trend
VLM4VLA : VLM scores don't predict control; frozen ViT β21~42 pp; action-supervised vision FT +18.1 pp β the one proven fix. 2026 splits into geometry add-ons vs dynamics encoders
βοΈ Approaches
action-align the VLM ViT (cheap, proven) Β· bolt-on metric 3D (spatial gains / calibration tax) Β· video-dynamics encoders (physics priors / loses VLM semantics)
β οΈ Open
Effects visible only under OOD stress; no standardized recipe
Which training framework / codebase? β modular vs pipeline; RL support is the new differentiator
π Trend
StarVLA (modular) vs TRI VLA Foundry (pipeline) β RSS 2026 adds the RL tier (RLux-VLA)
β οΈ Open
Nothing covers pretrainβSFTβRLβevaluation end-to-end; flagship recipes irreproducible in public stacks
β‘ Running & improving it
How do we make inference real-time? β from test-time patches to trained-in continuation
π Trend
RTC (test-time, 2025) β training-time RTC β Legato native continuation (beats RTC ~10%) + OAT anytime decoding (2026)
βοΈ Approaches
test-time guidance (drop-in / unstable on some models) Β· trained continuation (smooth / retrain) Β· anytime tokens (compute dial / AR-only) Β· async dual-rate (decoupled rates / complexity)
β οΈ Open
ms/chunk almost never published; robustness features raise the denoise-step bill (4β10)
How does a deployed policy improve from experience? π β production-proven at RSS 2026
π Trend
"Impractical" (2025) β log-prob wall cracked 3 ways β Ο*0.6/RECAP in production: 2Γ throughput, ~Β½ failures on laundry/boxes/espresso
βοΈ Approaches
advantage conditioning (no log-probs, eats corrections / coarse credit) Β· direct PG on flow (principled / machinery) Β· residual RL (safe / capped) Β· world-model RFT (no rollouts / WM fidelity)
β οΈ Open
Loops are domain-narrow; exploration safety procedural; forgetting is milder than feared (ICML Oral: pretrained VLAs resist it) but untested under repeated RL
How do we handle long-horizon tasks? β from bigger context to agentic decomposition
π Trend
In-policy memory modules (2025) β selective key-frame history + agentic planner/executor with notebook memory (RobotNav: EQA SOTA, β77% steps)
βοΈ Approaches
in-policy memory (fast / shortcut risk) Β· episodic retrieval (scales / retrieval-bound) Β· agentic notebook (auditable / planner latency)
β οΈ Open
Manipulation autonomy record is 11 stages / 2.5 min; no agentic-manipulation demo yet
Do world models help VLA? β no longer hypothetical: dynamics pre-training beats VLAs on hardware
π Trend
Serving role (2025) β challenger evidence at RSS 2026: LDA-1B +48% dexterous over Ο0.5; mimic-video 10Γ sample efficiency
βοΈ Approaches
data engine (mature / fidelity ceiling) Β· evaluator (promising / action-input gap β ICML's dWorldEval is the first crack) Β· VAM backbone (dynamics priors / loses VLM depth) Β· unified WM+policy (all data tiers / heaviest) Β· aux losses (free / weak)
β οΈ Open
Contact physics fidelity; VAM claims await matched-scale replication; 20B video models price out labs
Can human video substitute for robot data? π β yes; the field split on how
π Trend
Retargeting pipelines (2024β25) β three camps at RSS 2026: emergence (PI: co-train, ~2Γ above a diversity threshold) vs decoupling (Ξ¨β: 800 h + 30 h beats 10Γ corpora) vs synthesis (RobotManip: 24.8k h H2R)
βοΈ Approaches
co-train (no pipeline / threshold, gripper-only evidence) Β· staged decouple (data-efficient, humanoid-proven / 2-stage) Β· synthesize (unlimited scale / artifact ceiling)
β οΈ Open
Nobody has run both recipes on the same platform β the field's most valuable missing experiment
What data should we co-train on? π β the most settled question on this map
π Trend
Folklore β measurement: two LBM studies (89 policies, 58k+2,835 rollouts) + independent Qwen/VLM4VLA confirmations
β
Verdicts
VL + cross-embodiment data help cumulatively Β· discrete robot -action tokens don't (3Γ replicated; latent -action tokens as VLM supervision do β ICML Oral) Β· co-training must be joint , not sequential
β οΈ Open
Mixture ratios are art (Ξ»=0.1 vs 1.0, undiscussed); no per-sample attribution
How do we evaluate policies credibly? π β the in-distribution era is ending
π Trend
Triple indictment (from-scratch β pretrained in-distribution; LIBERO-X 39.4β8.2 collapse; weak sim-real correlation) β RSS 2026 infrastructure: PolaRiS real-to-sim with validated rank correlation
βοΈ Approaches
static sim (cheap / saturated) Β· perturbation pyramids (diagnostic / still sim) Β· real-to-sim (reality-anchored / contact physics) Β· rigorous real stats (ground truth / cost)
β οΈ Open
Contact-rich sim evaluation missing everywhere; OOD reporting still voluntary
How do we improve dexterous manipulation? β touch moved from observation to prediction
π Trend
RL still owns in-hand skills; IL changed paradigm at RSS 2026 β ViTacFormer forecasts future contact (+50%, 11-stage record); CGP commands contacts, not poses
βοΈ Approaches
sim-to-real RL (reactive / per-skill, tactile sim gap) Β· predictive visuo-tactile IL (breadth / rig cost) Β· dexterous VLAs (language / trails specialists)
β οΈ Open
No dexterous foundation model; precision insertion unsolved everywhere (Ξ¨β 2/10, screws 0/10)
How do we generalize across robots? β alignment is the precondition for data scaling
π Trend
Question flipped in 2026: RobotManip showed misaligned action spaces produce no scaling law ; camera-frame EEF scales + transfers zero-shot. RSS extended to hands (DexGrasp-Zero 85% zero-shot, One-Hand 81.9%)
βοΈ Approaches
canonical padded tensors (simple, proven / curated slots) Β· camera-frame delta EEF (best transfer / calibration tax) Β· morphology graphs/URDFs (anatomy-grounded / hands-only) Β· prompts + history (no arch change / weak alone)
β οΈ Open
Joint-space zero-shot <5%; Franka-class morphology gaps resist; gripperβhand untried
How do we run a VLA on a humanoid? β the triple-system recipe became the open reference
π Trend
RSS 2026 = humanoid loco-manipulation's coming-out; Ξ¨β 's VLM + MM-DiT + RL-lower-body triple system beats 10Γ-data baselines by +40 pp
βοΈ Approaches
whole-body end-to-end (expressive / unstable, data-hungry) Β· triple-system (stable, data-efficient / agility capped) Β· teleop rigs vs robot-free human interfaces (quality vs scale)
β οΈ Open
Per-task fine-tuning in every loop; no humanoid OOD protocol β evaluation lags arms by a generation
Newest first β full history in Changelog . Pre-publication preprint reviews live in Latest Papers .
β Topic reviews catalog β cross-paper topic reviews, lab programs, latest-paper reviews Β· Per-paper long-forms β single-paper deep-dives.
π 4. Conference Surveys
Venue
Year
Survey / Index
Character
RSS
2026
Survey π
Method frontier β the improvement loop, humanoids, hands, evaluation (210 papers)
ICRA
2026
Survey
Deployment & sensors at scale (728 in-scope papers, 12 topic pages)
ICML
2026
Index
99 manipulation papers + π
Top-20
CVPR
2026
Survey
Vision-first VLA / perception (also 2025 )
ICLR
2026
Survey
Architecture & representation (~211 papers)
NeurIPS
2025
Survey
Foundation-model training methods
CoRL
2025
Survey
Robot learning / benchmarks
IROS
2026
Survey π
Robustness/efficiency/deployment of learned & VLA policies (1,900+ papers; titles-level, pre-conf)
IROS
2025
Survey
Hardware, systems, deployment
π 5. Foundational References
OpenVLA (open 7B AR baseline) Β· ReKep (training-free keypoint constraints) Β· AgiBot World Colosseo (1M+ trajectory dataset) Β· RoboBrain 2.0 (embodied-reasoning VLM)
π Contributing: page templates & house rules β Maintenance . Summaries are compiled from public sources and are not substitutes for the original papers β verify numbers before citing.