Home - Heungwoo/research GitHub Wiki

πŸ€– VLA / Robot AI Research Review Wiki

A living research wiki for Vision-Language-Action models Β· robot manipulation Β· humanoid intelligence Β· tactile & force interaction Β· world models.

πŸ“„ 396 pages πŸ—“ 8 venues πŸ“– 16+ cross-paper reviews πŸ”„ updated 2026-07-25

New here? Start with one of these: 🧭 VLA Architectures β€” choose the model family & action decoder Β· πŸ“ˆ Ο€ Series Evolution β€” the production-VLA baseline lineage Β· πŸ† IROS 2026 Survey πŸ†• Β· RSS 2026 Survey β€” the latest venue surveys Β· πŸ†• Latest Papers β€” preprint tracker (pre-publication reviews) Β· πŸ•ΈοΈ Knowledge Graph β€” the visual map


πŸ—Ί 1. Research Decision Map

Pick your question β†’ open the fold-out for a 4-line orientation β†’ follow the πŸ“„ deep-dive link for the full survey (trend arc Β· approach taxonomy Β· limitations). Every deep-dive page carries a matching "State of the Field" section, updated Aug 2026.

πŸ— Building a policy

Which VLA architecture should we use? β€” flow matching dominates; ordered tokens revived AR

πŸ“„ Deep dive β†’ VLA Architectures

πŸ“ˆ Trend AR tokens (2024) β†’ flow-matching experts standard (2025) β†’ flow universal in flagships + AR revived via OAT's anytime prefix decoding (2026)
βš–οΈ Approaches AR discrete tokens (LLM-native / lossy, slow) Β· flow-matching expert (precise, production-proven / no native log-probs) Β· discrete diffusion in-VLM (parallel decode / little traction)
⚠️ Open No matched-scale comparison across the three families; latency rarely published
How should the VLM connect to the action expert? β€” evidence tilts to cross-attention, margins small

πŸ“„ Deep dive β†’ VLM↔Action Connection

πŸ“ˆ Trend 2025 four-way standoff β†’ first direct ablations in 2026: cross-attention wins in RobotManip (87.5 vs 87.0) and MM-DiT beats naive DiT in Ξ¨β‚€
βš–οΈ Approaches same-stack MoE+prefix-KV (tight / hard to retrofit) Β· cross-attention (decoupled, small experts / one-way) Β· concatenation (simplest / re-attends per step) Β· adaLN token (cheapest / bottleneck)
⚠️ Open ~0.5 pp margins on one benchmark; the two Qwen flagships disagree internally
How should attention / KV-cache be designed? β€” the runtime now drives the choice

πŸ“„ Deep dive β†’ VLA Attention

πŸ“ˆ Trend Design follows the runtime since 2025 (RTC/streaming/prefix-KV); 2026 adds linear-attention backbones (Qwen3.5) with zero control-quality ablation
βš–οΈ Approaches prefix-KV + expert branch (cache reuse) Β· full re-attention per chunk (freshness) Β· streaming-token designs
⚠️ Open Latency is the field's least-reported number; multi-view KV behavior is folklore
Should vision be built outside the VLM? β€” the ViT is the proven bottleneck; fix = action supervision

πŸ“„ Deep dive β†’ Independent Visual Representation

πŸ“ˆ Trend VLM4VLA: VLM scores don't predict control; frozen ViT βˆ’21~42 pp; action-supervised vision FT +18.1 pp β€” the one proven fix. 2026 splits into geometry add-ons vs dynamics encoders
βš–οΈ Approaches action-align the VLM ViT (cheap, proven) Β· bolt-on metric 3D (spatial gains / calibration tax) Β· video-dynamics encoders (physics priors / loses VLM semantics)
⚠️ Open Effects visible only under OOD stress; no standardized recipe
Which training framework / codebase? β€” modular vs pipeline; RL support is the new differentiator

πŸ“„ Deep dive β†’ VLA Training Frameworks

πŸ“ˆ Trend StarVLA (modular) vs TRI VLA Foundry (pipeline) β†’ RSS 2026 adds the RL tier (RLux-VLA)
⚠️ Open Nothing covers pretrainβ†’SFTβ†’RLβ†’evaluation end-to-end; flagship recipes irreproducible in public stacks

⚑ Running & improving it

How do we make inference real-time? β€” from test-time patches to trained-in continuation

πŸ“„ Deep dive β†’ Real-Time Execution

πŸ“ˆ Trend RTC (test-time, 2025) β†’ training-time RTC β†’ Legato native continuation (beats RTC ~10%) + OAT anytime decoding (2026)
βš–οΈ Approaches test-time guidance (drop-in / unstable on some models) Β· trained continuation (smooth / retrain) Β· anytime tokens (compute dial / AR-only) Β· async dual-rate (decoupled rates / complexity)
⚠️ Open ms/chunk almost never published; robustness features raise the denoise-step bill (4β†’10)
How does a deployed policy improve from experience? πŸ†• β€” production-proven at RSS 2026

πŸ“„ Deep dive β†’ RL for VLA

πŸ“ˆ Trend "Impractical" (2025) β†’ log-prob wall cracked 3 ways β†’ Ο€*0.6/RECAP in production: 2Γ— throughput, ~Β½ failures on laundry/boxes/espresso
βš–οΈ Approaches advantage conditioning (no log-probs, eats corrections / coarse credit) Β· direct PG on flow (principled / machinery) Β· residual RL (safe / capped) Β· world-model RFT (no rollouts / WM fidelity)
⚠️ Open Loops are domain-narrow; exploration safety procedural; forgetting is milder than feared (ICML Oral: pretrained VLAs resist it) but untested under repeated RL
How do we handle long-horizon tasks? β€” from bigger context to agentic decomposition

πŸ“„ Deep dive β†’ VLA Memory Β· System 0/1/2

πŸ“ˆ Trend In-policy memory modules (2025) β†’ selective key-frame history + agentic planner/executor with notebook memory (RobotNav: EQA SOTA, βˆ’77% steps)
βš–οΈ Approaches in-policy memory (fast / shortcut risk) Β· episodic retrieval (scales / retrieval-bound) Β· agentic notebook (auditable / planner latency)
⚠️ Open Manipulation autonomy record is 11 stages / 2.5 min; no agentic-manipulation demo yet
Do world models help VLA? β€” no longer hypothetical: dynamics pre-training beats VLAs on hardware

πŸ“„ Deep dive β†’ World Models Β· WAM vs VLA

πŸ“ˆ Trend Serving role (2025) β†’ challenger evidence at RSS 2026: LDA-1B +48% dexterous over Ο€0.5; mimic-video 10Γ— sample efficiency
βš–οΈ Approaches data engine (mature / fidelity ceiling) Β· evaluator (promising / action-input gap β€” ICML's dWorldEval is the first crack) Β· VAM backbone (dynamics priors / loses VLM depth) Β· unified WM+policy (all data tiers / heaviest) Β· aux losses (free / weak)
⚠️ Open Contact physics fidelity; VAM claims await matched-scale replication; 20B video models price out labs

πŸ“Š Data & evaluation

Can human video substitute for robot data? πŸ†• β€” yes; the field split on how

πŸ“„ Deep dive β†’ Human Video β†’ Robot Transfer

πŸ“ˆ Trend Retargeting pipelines (2024–25) β†’ three camps at RSS 2026: emergence (PI: co-train, ~2Γ— above a diversity threshold) vs decoupling (Ξ¨β‚€: 800 h + 30 h beats 10Γ— corpora) vs synthesis (RobotManip: 24.8k h H2R)
βš–οΈ Approaches co-train (no pipeline / threshold, gripper-only evidence) Β· staged decouple (data-efficient, humanoid-proven / 2-stage) Β· synthesize (unlimited scale / artifact ceiling)
⚠️ Open Nobody has run both recipes on the same platform β€” the field's most valuable missing experiment
What data should we co-train on? πŸ†• β€” the most settled question on this map

πŸ“„ Deep dive β†’ LBM Co-training (+ RSS 89-policy sequel)

πŸ“ˆ Trend Folklore β†’ measurement: two LBM studies (89 policies, 58k+2,835 rollouts) + independent Qwen/VLM4VLA confirmations
βœ… Verdicts VL + cross-embodiment data help cumulatively Β· discrete robot-action tokens don't (3Γ— replicated; latent-action tokens as VLM supervision do β€” ICML Oral) Β· co-training must be joint, not sequential
⚠️ Open Mixture ratios are art (λ=0.1 vs 1.0, undiscussed); no per-sample attribution
How do we evaluate policies credibly? πŸ†• β€” the in-distribution era is ending

πŸ“„ Deep dive β†’ VLA Evaluation

πŸ“ˆ Trend Triple indictment (from-scratch β‰ˆ pretrained in-distribution; LIBERO-X 39.4β†’8.2 collapse; weak sim-real correlation) β†’ RSS 2026 infrastructure: PolaRiS real-to-sim with validated rank correlation
βš–οΈ Approaches static sim (cheap / saturated) Β· perturbation pyramids (diagnostic / still sim) Β· real-to-sim (reality-anchored / contact physics) Β· rigorous real stats (ground truth / cost)
⚠️ Open Contact-rich sim evaluation missing everywhere; OOD reporting still voluntary

🦾 Embodiment

How do we improve dexterous manipulation? β€” touch moved from observation to prediction

πŸ“„ Deep dive β†’ Dexterous Manipulation Β· Tactile VLA

πŸ“ˆ Trend RL still owns in-hand skills; IL changed paradigm at RSS 2026 β€” ViTacFormer forecasts future contact (+50%, 11-stage record); CGP commands contacts, not poses
βš–οΈ Approaches sim-to-real RL (reactive / per-skill, tactile sim gap) Β· predictive visuo-tactile IL (breadth / rig cost) Β· dexterous VLAs (language / trails specialists)
⚠️ Open No dexterous foundation model; precision insertion unsolved everywhere (Ξ¨β‚€ 2/10, screws 0/10)
How do we generalize across robots? β€” alignment is the precondition for data scaling

πŸ“„ Deep dive β†’ Cross-Embodiment

πŸ“ˆ Trend Question flipped in 2026: RobotManip showed misaligned action spaces produce no scaling law; camera-frame EEF scales + transfers zero-shot. RSS extended to hands (DexGrasp-Zero 85% zero-shot, One-Hand 81.9%)
βš–οΈ Approaches canonical padded tensors (simple, proven / curated slots) Β· camera-frame delta EEF (best transfer / calibration tax) Β· morphology graphs/URDFs (anatomy-grounded / hands-only) Β· prompts + history (no arch change / weak alone)
⚠️ Open Joint-space zero-shot <5%; Franka-class morphology gaps resist; gripper↔hand untried
How do we run a VLA on a humanoid? β€” the triple-system recipe became the open reference

πŸ“„ Deep dive β†’ Humanoid VLA Β· System 0/1/2

πŸ“ˆ Trend RSS 2026 = humanoid loco-manipulation's coming-out; Ξ¨β‚€'s VLM + MM-DiT + RL-lower-body triple system beats 10Γ—-data baselines by +40 pp
βš–οΈ Approaches whole-body end-to-end (expressive / unstable, data-hungry) Β· triple-system (stable, data-efficient / agility capped) Β· teleop rigs vs robot-free human interfaces (quality vs scale)
⚠️ Open Per-task fine-tuning in every loop; no humanoid OOD protocol β€” evaluation lags arms by a generation

πŸ†• 2. Latest Updates

Newest first β€” full history in Changelog. Pre-publication preprint reviews live in Latest Papers.

Date Page What it adds
07-25 RSS 2026 survey + 15 paper pages Sydney; 210 papers, ~116 in scope; Ο€*0.6/RECAP flagship; figure-illustrated reviews
07-25 Ξ¨β‚€ (in-depth) Open humanoid foundation model β€” 800 h human video + 30 h robot data beats 10Γ— corpora
07-24 Qwen Team's VLA Program Cross-paper: VLM4VLA β†’ Qwen-VLA β†’ Qwen-Robot Suite (5 in-depth reviews)
07-24 Qwen-RobotManip Β· RobotNav Β· RobotWorld The Qwen-Robot Suite, deep-read
06-11 VLA Training Frameworks StarVLA vs TRI VLA Foundry
06-10 RoboMME (in-depth) Memory-implementation breakdown

πŸ“– 3. In-Depth Reviews

β†’ Topic reviews catalog β€” cross-paper topic reviews, lab programs, latest-paper reviews Β· Per-paper long-forms β€” single-paper deep-dives.

Most-used entries
Topic reviews VLA Architectures Β· VLA Hybrid Architectures πŸ†• Β· Multi-Task VLA πŸ†• Β· VLM↔Action Β· RL for VLA Β· World Models Β· Dexterous Β· Dex-Hand Data Pyramid πŸ†• Β· Tactile Β· Cross-Embodiment Β· Single-Checkpoint Multi-Robot πŸ†• Β· Egocentric Video Pre-Training πŸ†• Β· Humanoid Β· Memory Β· In-Context Imitation πŸ†•
Lab programs Ο€ series Β· GR00T N1β†’N1.7 Β· RoboTTT (context scaling) πŸ†• Β· Qwen VLA program Β· Ξ¨β‚€ Β· DreamZero πŸ†•
ML foundations ML hub Β· Attention Variants Β· Normalization

πŸ—“ 4. Conference Surveys

Venue Year Survey / Index Character
RSS 2026 Survey πŸ†• Method frontier β€” the improvement loop, humanoids, hands, evaluation (210 papers)
ICRA 2026 Survey Deployment & sensors at scale (728 in-scope papers, 12 topic pages)
ICML 2026 Index 99 manipulation papers + πŸ… Top-20
CVPR 2026 Survey Vision-first VLA / perception (also 2025)
ICLR 2026 Survey Architecture & representation (~211 papers)
NeurIPS 2025 Survey Foundation-model training methods
CoRL 2026 Survey πŸ†• Robot learning (687 accepted; preliminary β€” official list pending)
CoRL 2025 Survey Robot learning / benchmarks
IROS 2026 Survey πŸ†• Robustness/efficiency/deployment of learned & VLA policies (1,900+ papers; titles-level, pre-conf)
IROS 2025 Survey Hardware, systems, deployment

πŸ“Œ 5. Foundational References

OpenVLA (open 7B AR baseline) Β· ReKep (training-free keypoint constraints) Β· AgiBot World Colosseo (1M+ trajectory dataset) Β· RoboBrain 2.0 (embodied-reasoning VLM)


πŸ›  Contributing: page templates & house rules β†’ Maintenance. Summaries are compiled from public sources and are not substitutes for the original papers β€” verify numbers before citing.

⚠️ **GitHub.com Fallback** ⚠️