Reviews - Heungwoo/research GitHub Wiki
π In-Depth Reviews β full catalog
Every cross-paper topic review, lab/series program review, and latest-paper review in this wiki. Per-paper single-paper long-forms live on a companion page: Per-Paper Long-Forms. The sidebar links only the most-used entries.
β Back to [Home]] Β· visual map: [Knowledge Graph
1. Cross-paper topic reviews
Taxonomy + comparison tables + decision guide each.
VLA core
- VLA Architectures β 13 action-decoder categories; the canonical "what architecture?" review (sub-pages: category details Β· paper table)
- VLMβAction Connection β 7 ways the VLM conditions the action expert
- VLA Attention β per-family attention masks (Ο Β· GR00T Β· StarVLA Β· Qwen-VL)
- VLA Memory β 6 memory architectures + decision guide
- In-Context Imitation & Demo-Following π β watch a demo, reproduce it (no per-task FT): cross-attention Β· recurrent Β· fast-weight/TTT Β· retrieval Β· token-ICL, mapped to RoboMME's imitation suite
- Goal-Image Conditioning β goal/subgoal-image VLAs vs Ο0.7
- Independent Visual Representation β when vision is built outside the VLM
- VLA Hybrid Architectures π β the WAM+VLA convergence; three-expert MoT with vision as a separate tower (Motus Β· BagelVLA Β· HALO Β· BAGEL) + a design guide
- Multi-Task VLA π β why one policy fails across many tasks (negative transfer, non-mergeability, forgetting) + the fix landscape (MergeVLA merging Β· MoE Β· gradient Β· instruction grounding Β· continual)
- HiMoE-VLA π β hierarchical depth-wise MoE (action-space β embodiment) beats negative transfer
- DyGRO-VLA π β cross-task RL fine-tuning without forgetting (protect shared latent + grouped RL residuals)
- Motus π β unified scheduled MoT (understanding+video-gen+action, 8B, open)
- HALO π β three-expert EM-CoT VLA (thinkβimagineβact); ablates the vision tower's value
- BAGEL π β the base multimodal MoT recipe (VAE+ViT dual encoders) robot VLAs inherit
Learning & training
- RL for VLA β residual Β· RL-token Β· outcome-conditioned Β· world-model RFT (RSS 2026: + advantage conditioning)
- VLA Training Frameworks β StarVLA (Lego-modular) vs TRI VLA Foundry
- LBM Co-training Study β TRI's co-training evidence base (sequel: RSS 89-policy study)
- Human Video β Robot Transfer π β the emergence / decoupling / synthesis fork
- Egocentric Video for Pre-Training π β how label-free first-person video becomes a pretraining signal (pseudo-actions Β· latent actions Β· world models Β· retargeting) + datasets + scaling laws
- Real-Time Execution π β chunking, continuation, anytime decoding, async systems
- RoboTTT (context scaling) π β NVIDIA GEAR: 8K-timestep context via Test-Time-Training fast weights in GR00T N1.7's DiT (constant latency)
- Stellar VLA (continual imitation learning) π β Dirichlet-Process self-evolving knowledge space + knowledge-routed MoE (preprint)
Evaluation methodology
- VLA Evaluation π β the in-distribution indictment + the 2026 toolkit (perturbation pyramids, real-to-sim, statistics)
Robot capability & sensing
- Dexterous Manipulation β 70+ papers; "RL is still the dex core"
- Dexterous-Hand Data Pyramid π β data types (web-video β glove β retarget β sim β teleop) + one approachΓdata matrix + current/future insight
- Do As I Do π β device-free everyday video β dexterous data via 4D reconstruction + dynamics-aware retargeting
- AnyDexRT π β calibration-free, cross-hand humanβrobot retargeting (the L4 bridge, 7 hands)
- YUBI π β handheld finger-driven gripper, deploy w/o retargeting; 8,434 h / 1.20M ep bimanual dataset
- DexEXO π β wearability-first exoskeleton, visual-match to a 6-DoF hand; beats DexUMI/teleop on contact-rich tasks
- T-Rex π β variable-rate MoT with a fast tactile expert; reactive force control, +30pts over EgoScale on delicate tasks
- RLDX-1 π β RLWRLD's dexterity-first foundation model (MSAT 4-stream); human-hand-first data (vendor claims)
- Genesis GENE-26.5 π β Genesis AI's glove-first dexterous FM; 1:1:1 tactile glove, <1h robot fine-tune (vendor claims)
- Tactile VLA β how touch enters the policy Γ sensor hardware Γ trends
- Cross-Embodiment β one policy, many bodies (training-data cut)
- Single-Checkpoint Multi-Robot Deployment π β one frozen checkpoint controlling many robots at inference (routed seen-robot vs zero-shot-to-unseen: LAP Β· Green-VLA Β· Gemini Robotics 1.5 Β· RT-X Β· CrossFormer Β· Ο0.7)
- Humanoid VLA β whole-body & bipedal loco-manipulation
- System 0 / 1 / 2 β the cognitive-tier framing for humanoid stacks
World models, evaluation & robustness
- World Models β what a world model predicts Γ how robotics uses it
- WAM vs VLA Robustness β first controlled world-model-vs-VLA benchmark
- NVIDIA WAM thesis + Cosmos 3 β the "imagineβact" thesis deep-dive
- DreamZero (WAM-as-policy) π β NVIDIA's 14B video-diffusion World Action Model; >2Γ over VLAs, real-time 7 Hz
- Ο-0 (humanoid latent WAM) π β reconstruction-free latent future prediction for concurrent whole-body loco-manipulation (preprint)
- DYNA-2 (industry WAM) π β Dyna Robotics' human-video-only world-action model; claimed 1M-hour scaling law (company release β vendor claims)
ML foundations
1b. Latest-paper reviews (preprints)
Reviewed ahead of venue publication β see Latest Papers.
- Ο-0 (humanoid WAM) π β latent-predictive whole-body loco-manipulation
- Stellar VLA (continual learning) π β Dirichlet-Process knowledge space for CIL
- DYNA-2 (WAM, human-video scaling) π β Dyna Robotics' world-action model release (company claims)
- Motus (unified latent-action WAM) π β one scheduled 8B MoT = world model / VLA / IDM / video gen (open weights)
- Cortex 2.0 (foresight-planning hybrid) π β Sereact's world-model planner scores imagined futures (deployment-reported)
- Being-H0.7 (latent world-action model) π β posterior/prior latent reasoning, reactive at inference
2. Lab & series programs
- Ο Series Evolution β Ο0 β Ο0.5 β Ο0.6 β Ο*0.6/RECAP β Ο0.7; companion baselines: Ο0.6 Β· Ο0.7 Β· Ο*0.6 + RECAP; long-forms: Ο0.6 Β· Ο0.7
- GR00T N1 β N1.7 β NVIDIA's open humanoid line
- Qwen Team's VLA Program β VLM4VLA β Qwen-VLA β Qwen-Robot Suite; constituents: Qwen-VLA Β· Qwen-RobotManip Β· Qwen-RobotNav Β· Qwen-RobotWorld Β· VLM4VLA
- Ξ¨β humanoid foundation model β USC PSI Lab Γ NVIDIA open humanoid stack (RSS 2026)
3. Per-paper long-forms β moved to a dedicated page
β Per-Paper Long-Forms β single-paper deep-dives (architecture/runtime Β· data/training/eval Β· world-models/tactile Β· hybrid-MoT Β· dexterous-hand data Β· multi-task/in-context Β· IROS 2026 full-paper analyses Β· RSS 2026 per-paper). (Split out so this catalog renders quickly.)
When adding a review: topic reviews go here, per-paper long-forms go in Reviews-Per-Paper; note it in Changelog; the sidebar stays top-level-only.
β Back to Home