Reviews - Heungwoo/research GitHub Wiki

πŸ“– In-Depth Reviews β€” full catalog

Every cross-paper topic review, lab/series program review, and latest-paper review in this wiki. Per-paper single-paper long-forms live on a companion page: Per-Paper Long-Forms. The sidebar links only the most-used entries.

← Back to [Home]] Β· visual map: [Knowledge Graph


1. Cross-paper topic reviews

Taxonomy + comparison tables + decision guide each.

VLA core

  • VLA Architectures β€” 13 action-decoder categories; the canonical "what architecture?" review (sub-pages: category details Β· paper table)
  • VLM↔Action Connection β€” 7 ways the VLM conditions the action expert
  • VLA Attention β€” per-family attention masks (Ο€ Β· GR00T Β· StarVLA Β· Qwen-VL)
  • VLA Memory β€” 6 memory architectures + decision guide
  • In-Context Imitation & Demo-Following πŸ†• β€” watch a demo, reproduce it (no per-task FT): cross-attention Β· recurrent Β· fast-weight/TTT Β· retrieval Β· token-ICL, mapped to RoboMME's imitation suite
  • Goal-Image Conditioning β€” goal/subgoal-image VLAs vs Ο€0.7
  • Independent Visual Representation β€” when vision is built outside the VLM
  • VLA Hybrid Architectures πŸ†• β€” the WAM+VLA convergence; three-expert MoT with vision as a separate tower (Motus Β· BagelVLA Β· HALO Β· BAGEL) + a design guide
  • Multi-Task VLA πŸ†• β€” why one policy fails across many tasks (negative transfer, non-mergeability, forgetting) + the fix landscape (MergeVLA merging Β· MoE Β· gradient Β· instruction grounding Β· continual)
    • HiMoE-VLA πŸ†• β€” hierarchical depth-wise MoE (action-space β†’ embodiment) beats negative transfer
    • DyGRO-VLA πŸ†• β€” cross-task RL fine-tuning without forgetting (protect shared latent + grouped RL residuals)
    • Motus πŸ†• β€” unified scheduled MoT (understanding+video-gen+action, 8B, open)
    • HALO πŸ†• β€” three-expert EM-CoT VLA (thinkβ†’imagineβ†’act); ablates the vision tower's value
    • BAGEL πŸ†• β€” the base multimodal MoT recipe (VAE+ViT dual encoders) robot VLAs inherit

Learning & training

Evaluation methodology

  • VLA Evaluation πŸ†• β€” the in-distribution indictment + the 2026 toolkit (perturbation pyramids, real-to-sim, statistics)

Robot capability & sensing

  • Dexterous Manipulation β€” 70+ papers; "RL is still the dex core"
  • Dexterous-Hand Data Pyramid πŸ†• β€” data types (web-video β†’ glove β†’ retarget β†’ sim β†’ teleop) + one approachΓ—data matrix + current/future insight
    • Do As I Do πŸ†• β€” device-free everyday video β†’ dexterous data via 4D reconstruction + dynamics-aware retargeting
    • AnyDexRT πŸ†• β€” calibration-free, cross-hand humanβ†’robot retargeting (the L4 bridge, 7 hands)
    • YUBI πŸ†• β€” handheld finger-driven gripper, deploy w/o retargeting; 8,434 h / 1.20M ep bimanual dataset
    • DexEXO πŸ†• β€” wearability-first exoskeleton, visual-match to a 6-DoF hand; beats DexUMI/teleop on contact-rich tasks
    • T-Rex πŸ†• β€” variable-rate MoT with a fast tactile expert; reactive force control, +30pts over EgoScale on delicate tasks
    • RLDX-1 πŸ†• β€” RLWRLD's dexterity-first foundation model (MSAT 4-stream); human-hand-first data (vendor claims)
    • Genesis GENE-26.5 πŸ†• β€” Genesis AI's glove-first dexterous FM; 1:1:1 tactile glove, <1h robot fine-tune (vendor claims)
  • Tactile VLA β€” how touch enters the policy Γ— sensor hardware Γ— trends
  • Cross-Embodiment β€” one policy, many bodies (training-data cut)
  • Single-Checkpoint Multi-Robot Deployment πŸ†• β€” one frozen checkpoint controlling many robots at inference (routed seen-robot vs zero-shot-to-unseen: LAP Β· Green-VLA Β· Gemini Robotics 1.5 Β· RT-X Β· CrossFormer Β· Ο€0.7)
  • Humanoid VLA β€” whole-body & bipedal loco-manipulation
  • System 0 / 1 / 2 β€” the cognitive-tier framing for humanoid stacks

World models, evaluation & robustness

ML foundations


1b. Latest-paper reviews (preprints)

Reviewed ahead of venue publication β€” see Latest Papers.

2. Lab & series programs


3. Per-paper long-forms β†’ moved to a dedicated page

β†’ Per-Paper Long-Forms β€” single-paper deep-dives (architecture/runtime Β· data/training/eval Β· world-models/tactile Β· hybrid-MoT Β· dexterous-hand data Β· multi-task/in-context Β· IROS 2026 full-paper analyses Β· RSS 2026 per-paper). (Split out so this catalog renders quickly.)


When adding a review: topic reviews go here, per-paper long-forms go in Reviews-Per-Paper; note it in Changelog; the sidebar stays top-level-only.

← Back to Home