Review Humanoid VLA - Heungwoo/research GitHub Wiki

In-Depth Review — VLA for Humanoids (whole-body & bipedal loco-manipulation)

Compiled June 2026 · Focus: what makes a Vision-Language-Action model for a humanoid technically different from a tabletop-arm VLA — specifically (1) mobility / bipedal balance during manipulation, (2) the high-dimensional whole-body action space, and (3) data collection when you cannot simply teleoperate a balancing biped.

Companion reviews: System 0/1/2 architectures · GR00T series · Dexterous Manipulation · Cross-Embodiment · VLA Architectures · DuoCore-FS.


1. TL;DR — humanoid VLA is a layered problem, not a bigger arm

A tabletop VLA maps (image, instruction) → end-effector action. A humanoid cannot be modeled that way, for three reasons that organize this whole review:

  1. The robot can fall over. Every action is conditioned on keeping a ~19–29-DoF underactuated biped balanced. Almost no system lets the VLA emit balance-critical torques directly — balance is delegated to a separate RL whole-body controller (WBC) and the VLA talks to it through a latent or a command.
  2. The action space is 2–5× larger and heterogeneous. Arms + hands + waist + neck + legs, mixing position, force, and locomotion commands at different rates. The field's answers are latent action vocabularies, dual-system VLM→fast-policy splits, and stream-specific tokenization — not one monolithic head.
  3. You cannot cheaply teleoperate a walking humanoid. So the data engine shifts off teleop toward human video → reconstruction → retargeting, motion-capture → RL-tracking → distillation, and sim-RL teacher → vision student. Teleop survives mostly for stationary upper-body work (Helix, GR00T's real layer).

The single most important structural fact: "humanoid VLA" today is a stack, cleanly described by the System 0/1/2 framing — a slow VLM (System 2, 1–10 Hz), a fast visuomotor policy (System 1, 50–200 Hz), and a reflex/balance layer (System 0, 500 Hz–1 kHz). The genuinely-new research is in how the VLA hands off to the balance layer and where the training data comes from.

Headline caveat — read before believing any "whole-body" claim: several flagship "humanoid VLAs" are upper-body only and never walk — Figure Helix (35-DoF, upper body, no paper, vendor numbers only) and GR00T's public GR-1 deployment (stationary bimanual). True bipedal loco-manipulation VLAs — where the policy commands walking/balance — remain rare and mostly simulation-trained: LeVERB, WholeBodyVLA, HumanVLA, Humanoid-VLA.


2. Why a humanoid breaks the tabletop-VLA assumptions

Tabletop-VLA assumption Why a humanoid violates it Field's response
The base is fixed; actions don't threaten stability Underactuated biped; any arm motion shifts CoM and can topple it Delegate balance to an RL WBC (System 0/1); VLA emits latent/command, not torque
≤7–14 DoF, one action modality (EE pose) 19–36 DoF spanning legs/waist/arms/hands/neck; mixes position + force + gait Latent action vocab · dual-system split · stream-specific tokenization · hierarchical value decomposition
Teleop scales data linearly A walking biped is unsafe/awkward to teleoperate; balance data isn't demonstrable Human-video reconstruction · mocap→RL→distill · sim teacher→vision student
One control rate (~10–50 Hz) Reasoning (1–10 Hz) and balance (≥500 Hz) differ by ~100× Asynchronous fast-slow inference; per-rate heads
One embodiment per policy Many humanoid morphologies (G1, H1, GR-1, Atlas, AgiBot, Astribot) Per-embodiment MLP projection to a shared latent; cross-embodiment pyramids

This is why humanoid VLA looks like "a VLM bolted onto a learned whole-body controller" rather than "scale imitation until it works."


3. The landscape — three layers and four families

The three-layer stack — a slow VLM hands off to a fast policy, which hands off to a balance layer:

flowchart TB
  VLM["System 2 · slow VLM · 1–10 Hz<br/>latent / subtask / motion code"]
  POL["System 1 · fast policy · 50–200 Hz<br/>diffusion / flow / AR head"]
  WBC["System 0 · reflex & balance · 500 Hz–1 kHz<br/>RL whole-body controller"]
  ROB["humanoid joints · 19–36 DoF"]
  VLM --> POL --> WBC --> ROB
Loading

The four families — which approach fits your constraint:

flowchart TB
  Q{Which family?}
  Q -- learned-latent VLA over an RL WBC --> F1[F1. Latent-vocabulary whole-body VLA<br/>LeVERB · WholeBodyVLA]
  Q -- VLM then fast-policy split --> F2[F2. Dual-system manipulation VLA<br/>GR00T · Helix · DuoCore-FS · Galaxea]
  Q -- language to motion / limited vision --> F3[F3. Language-to-motion controllers<br/>Lang-To-Loco · BFM-Zero · Humanoid-GPT]
  Q -- sim or human-video data engine --> F4[F4. Data & sim-to-real engines<br/>VideoMimic · VIRAL · EgoScale · Open-Sim-to-Real]
Loading
  • F1 — Latent-vocabulary whole-body VLA. The VLA emits a learned latent action token; an RL WBC decodes it into balanced whole-body joint targets. The cleanest answer to all three axes at once. (LeVERB, WholeBodyVLA.)
  • F2 — Dual-system manipulation VLA. Slow VLM + fast policy, heterogeneous DoF handled by per-embodiment MLPs. Strong on manipulation; balance is usually out of scope (upper-body or mobile-base). (GR00T N1–N1.7, Helix, DuoCore-FS, Galaxea G0, Fast-in-Slow.)
  • F3 — Language-to-motion / promptable controllers. Bridge language (sometimes not vision) to whole-body motion via latents; these are often the System 0/1 substrate a full VLA sits on. (Lang-To-Loco/RoboGhost, BFM-Zero, Humanoid-GPT, HWC-Loco, HVD, LIFT, UniFP.)
  • F4 — Data & sim-to-real engines. Not policies per se but the data-collection machinery that makes the others trainable. (VideoMimic, VIRAL, Open-Sim-to-Real, Gallant, EgoVLA, EgoScale.)

4. Axis 1 — Mobility & bipedal balance (the differentiator)

The defining humanoid problem. Three sub-patterns dominate:

4.1 Delegate balance to an RL whole-body controller (near-universal)

The VLA almost never outputs balance torques. Instead a low-level RL WBC owns stability and the VLA steers it:

  • LeVERB (arXiv 2506.13751, Berkeley): a System-1 RL WBC produces dynamics-feasible walking/turning/sitting from a latent instruction; the System-2 vision-language policy never touches joint torques.
  • WholeBodyVLA (ICLR 2026, 2512.11047, AgiBot X2): a Loco-Manipulation-Oriented (LMO) RL policy for advancing / turning / squatting — "manipulation-aware locomotion" that moves the base to serve the manipulation goal. This is the explicit bipedal-loco-manip contribution.
  • HumanVLA (NeurIPS 2024, 2406.19972): a state-based teacher trained with goal-conditioned RL + Adversarial Motion Prior (AMP) learns balanced walk-and-carry, then distilled into a VLA student.
  • BFM-Zero (ICLR 2026, 2511.04131, Unitree G1, 29-DoF): a promptable Forward-Backward foundation controller with disturbance rejection (recovers from kicks/pushes never trained on) — the substrate a VLA prompts.

4.2 Make locomotion robust and constraint-aware (the controller research)

These are not VLAs but the balance layer the field depends on:

  • HWC-Loco (ICLR 2026, H1/G1, 19-DoF): 3-stage hierarchy with a ZMP-feasibility constraint and robust optimization over an uncertainty set; survives ≤200 N pushes, 15 cm stairs, 20° slopes. Learned ZMP-threshold switching beats fixed thresholds (195 vs 459 switches).
  • UniFP (CoRL 2025 Best Paper, 2505.20829): a single unified policy for position and force in legged loco-manipulation — infers external force from proprioception (no F/T sensor), no hand-engineered mode switching; +39.5% on contact-rich tasks (wiping, drawers).
  • Gallant (CVPR 2026, 2511.14625): 3-D voxel-grid terrain perception (overhead + lateral constraints, not just a 2-D heightmap) for navigating pipes/shelves/stairs; near-100% staircase success.

4.3 The frequency stack (why this is hard in real time)

Reasoning, control, and balance run ~100× apart: VLM 1–10 Hz → fast policy 50–200 Hz → balance 500 Hz–1 kHz. Industrial systems make this explicit — Figure's reported System-0 is a 1 kHz neural prior trained on >1,000 h of human motion capture that replaced ~109k lines of C++ balance code (per Review-System-0-1-2). The async-inference papers (DuoCore-FS 2.6×, Fast-in-Slow 117.7 Hz) exist precisely to keep the fast layer fed while the slow VLM lags.

Reality check: the most-publicized "humanoid" VLAs sidestep Axis 1 entirely. Helix is upper-body only (no walking/balance claim). DuoCore-FS explicitly omits a System-0 balance layer (quasi-static kiosk). GR00T N1's GR-1 deployment is stationary. Bipedal balance during VLA-driven manipulation is demonstrated by a short list — LeVERB, WholeBodyVLA, HumanVLA — and mostly in simulation.


5. Axis 2 — The high-dimensional whole-body action space

DoF roughly doubles-to-quintuples vs a 7-DoF arm: H1 ~19, G1 ~23–29, Astribot S1 25, Figure 35 (upper only), Dexora 36 (bimanual+hands). Five distinct strategies:

Strategy Mechanism Representative Why it helps
A. Latent action vocabulary VLA emits a learned latent token; RL WBC decodes to whole-body joints LeVERB, WholeBodyVLA Decouples semantics from the 30-DoF control problem; no hand-crafted action primitives
B. Dual-system + per-embodiment projection Slow VLM → shared latent → per-embodiment MLP encoders/decoders → fast head GR00T N1–N1.7, Helix, Galaxea G0, Fast-in-Slow One model spans many morphologies; no explicit upper/lower split needed
C. Stream-specific tokenization Residual-VQ-VAE with parallel codebooks per stream (position / SO(3) / gripper) DuoCore-FS (29-dim → 36 fixed tokens, 3.4× more compact than FAST) Fixed-length tokens make AR decoding of whole-body actions tractable and fast
D. Hierarchical value decomposition Decompose the value function along kinematic structure (not the policy) HVD (WB-50 dataset) Credit assignment in high-DoF offline RL while keeping one unified policy
E. Behavioral latent / FB representation Map 29-DoF control into a low-dim latent z (e.g. R²⁵⁶); plan/track in latent BFM-Zero, Lang-To-Loco (64-D motion latent) Few-shot CEM/MPC and language-conditioning become low-dim problems

Notes that matter:

  • Explicit upper/lower-body decomposition is rare. Most systems use a shared latent and let projection layers sort out morphology. WholeBodyVLA's dual-stream head (separate arm-joint vs locomotion-command outputs) is one of the few explicit splits.
  • Frequency mismatch is part of the action-space problem. Helix runs S1 at 200 Hz (80M params) under a 7–9 Hz 7B VLM; FiS reuses only the last 2 of 32 LLaMA blocks at high rate; DuoCore-FS runs an async 25–30 Hz fast head under a 1–3 Hz slow head.
  • Dexterity rides on top. Bimanual 12-DoF hands (EgoVLA's Inspire hands, Dexora's XHAND) compound the DoF count — see Dexterous Manipulation for the hand-specific story; the whole-body papers mostly treat hands as another projected stream.

6. Axis 3 — Data collection (the real bottleneck)

You cannot teleoperate a balancing biped at scale, so humanoid VLA data comes from four routes. Retargeting (human morphology → robot URDF under joint limits) is the recurring crux.

6.1 Human video → reconstruction → retargeting

The dominant academic route — cheap monocular/egocentric human video lifted to robot actions:

  • VideoMimic (CoRL 2025 Best Student Paper, 2505.03729, G1): everyday RGB video → 4-D reconstruction (SMPL mesh + scene, metric scale via joint Levenberg-Marquardt optimization) → retarget → 4-stage sim (mocap pretrain → scene-conditioned tracking → DAgger distill → under-conditioned RL). Removing the mocap-pretrain stage breaks it.
  • EgoVLA (CVPR 2026, 2507.12440, H1 + 2×12-DoF Inspire): ~500K egocentric image-action pairs (HOI4D/HOT3D/HoloAssist/TACO), action = wrist SE(3) + MANO hand params, deployed via IK + hand retargeting. Pretraining lifts long-horizon success 2.22% → 45.93% vs ACT. Zero-shot (no robot finetune) = 0% — human video alone is insufficient.
  • EgoScale (arXiv 2602.16710): 20,854 h of action-labeled egocentric video (monocular SLAM + 21-keypoint hand pose + CasADi/IPOPT retargeting) — ~20× prior ego-video efforts; feeds GR00T N1.7 with a log-linear scaling law L_val = 0.024 − 0.003·ln(D) (R²=0.998).
  • WholeBodyVLA learns a Latent Action Model from action-free egocentric video — no demonstration labels — then a thin teleop layer grounds it.
  • Humanoid-VLA (arXiv 2502.14795): language-motion pre-alignment → egocentric video-conditioned PEFT → self-supervised pseudo-annotation of unlabeled video.

6.2 Motion capture → RL tracking experts → distillation

For locomotion/whole-body motion (where balance, not semantics, is the target):

  • Humanoid-GPT (CVPR 2026, G1): 2-billion-frame retargeted mocap (AMASS + LAFAN1 + Motion-X++ + MotionMillion + in-house), Harmonic-Motion-Embedding diversity clustering → per-cluster RL experts → DAgger into one causal GPT; zero-shot motion tracking at <1.5 ms (TensorRT).
  • BFM-Zero: LAFAN1 (40 motions retargeted to G1), EMD-prioritized sampling, 192 M env-steps off-policy.
  • Lang-To-Loco (RoboGhost): MotionMillion (50,378 seqs → 3,261 stable after filtering); language → 64-D motion latent → diffusion student; retargeting-free.

6.3 Sim-RL teacher → vision student (the sim-to-real engine)

Privileged-state RL in sim, distilled to an RGB policy that transfers zero-shot:

  • VIRAL (CVPR 2026, 2511.15200, NVIDIA/CMU/Berkeley, G1): privileged teacher → RGB student via large-scale tiled-rendering DAgger; compute is decisive (up to 64 GPUs; low-compute runs fail); 54 zero-shot real cycles.
  • Open-Sim-to-Real (CVPR 2026, 2512.01061): staged-reset PPO teacher → RGB student → GRPO finetune for partial observability; humanoid door-opening 31.7% faster than human teleoperators.
  • LeVERB and HumanVLA are entirely synthetic (rendered kinematic demos + RL); no teleop, no real video.

6.4 Teleoperation (the industrial route, survives for stationary/upper-body)

  • Figure Helix: ~500 h multi-robot multi-operator teleop (vendor figure; no paper).
  • GR00T: the "data pyramid" — 88 h GR-1 teleop (VIVE + Xsens) at N1, scaling to thousands of hours of real teleop by N1.6 (YAM/AGIBot/G1). The version-over-version story is which layer gets scaled: N1→N1.5 synthetic (DreamGen) · N1.5→N1.6 real-teleop · N1.6→N1.7 human-video (EgoScale).
  • Dexora: hybrid exoskeleton backpack + Apple Vision Pro teleop (gross arm + markerless fingers); 100 K sim + 10 K real episodes.
  • DuoCore-FS: just 10.22 h teleop, single kiosk task — efficient but narrow.

The data takeaway: academic whole-body VLAs lean synthetic / human-video; industrial manipulation VLAs lean teleop; GR00T is the only line that fuses all three at scale. Retargeting quality (metric-scale reconstruction, IK under joint limits) is the silent determinant of whether human-video data transfers.


7. Per-venue trends (2025–2026)

  • CoRL 2025 — human-video-as-cross-embodiment-source goes mainstream. VideoMimic (Best Student Paper) and UniFP (Best Paper) frame the year: monocular human video → balanced humanoid skills, and unified position-force loco-manipulation.
  • NeurIPS 2025 — async dual-system inference. Fast-in-Slow embeds the fast policy inside the VLM's last blocks (117.7 Hz) — the architectural enabler for running a slow reasoner over a fast humanoid.
  • ICLR 2026 — the whole-body controller substrate matures. A dense cluster of non-VLA controllers a VLA sits on: BFM-Zero (promptable FB foundation model), WholeBodyVLA (the standout true loco-manip VLA), Lang-To-Loco, HWC-Loco, HVD, LIFT.
  • CVPR 2026 — the sim-to-real data engine. VIRAL, Open-Sim-to-Real, Gallant, Humanoid-GPT, EgoVLA — the vision venue owns how you generate humanoid training signal at scale (teacher-student, voxel terrain, mocap pretraining, ego-video).
  • Industry (no papers) — dual-system + teleop. Figure Helix and NVIDIA GR00T define the deployed template; Helix is upper-body, GR00T is cross-embodiment with a real-teleop core.
  • IROS 2026 🆕 — bimanual coordination structure + whole-body unification. The frontier is coordination architecture, not a second arm: 3D FlowMatch Actor (CMU/NVIDIA — one 3D policy for single and dual-arm, +41.4% PerAct2, beats 1000×-larger models) and EquiBim (bilateral symmetry-equivariance as an inductive bias). Whole-body: ULTRA (unified multimodal whole-body loco-manip), CEER (compliant EE+root unified interface), OmniDP (beyond-FOV omnidirectional 3D perception), DreamMimic (loco-manip via a world model), and a MoE-VLA for humanoid loco-manip (Responsibility-Induced Specialized Experts — the multi-task interference fix reaching whole-body). Data is the bottleneck → RL-based bimanual data generation + physics-informed retargeting (SPIDER). Context: IROS 2026 survey §5.3.

8. Comparison table

System Venue arXiv/ID Platform · DoF Axis 1 — Mobility/Balance Axis 2 — Action space Axis 3 — Data
LeVERB arXiv Jun 2025 2506.13751 G1 · ~23–29 RL WBC decodes latent → walk/turn/sit Latent action vocabulary (A) Synthetic only; LeVERB-Bench 58.5%
WholeBodyVLA ICLR 2026 2512.11047 AgiBot X2 · bipedal LMO RL policy (advance/turn/squat) Unified latent + dual-stream head (A) Action-free ego video + thin teleop; +21.3%
HumanVLA NeurIPS 2024 2406.19972 Sim humanoid AMP teacher → balanced walk-carry Single whole-body policy Sim teacher→VLA student; Human-in-the-Room set
Humanoid-VLA arXiv Feb 2025 2502.14795 Humanoid (unspec) Whole-body control backbone — Language-motion align + ego video + pseudo-labels
GR00T N1–N1.7 NVIDIA 2503.14734 (N1) GR-1/G1/Atlas… N1 manipulation-only; G1 WBC in later code Dual-system DiT + per-embodiment MLP (B) Data pyramid (teleop+sim+ego); N1.7 = 20,854 h EgoScale
Figure Helix blog only no paper Figure 02 · 35 (upper) None (upper-body only) Dual-system 7B@7–9 Hz + 80M@200 Hz (B) ~500 h teleop (vendor)
DuoCore-FS arXiv Dec 2025 2512.20188 Astribot S1 · 25 Omitted (quasi-static) RVQ stream tokenizer (C); async 2.6× 10.22 h teleop, 1 task
Galaxea G0 arXiv Sep 2025 2509.00576 Mobile base Mobile (wheeled, not bipedal) Dual-system planner+executor (B) Open dataset; 3-stage curriculum
BFM-Zero ICLR 2026 2511.04131 G1 · 29 Promptable FB controller; push recovery FB latent z∈R²⁵⁶ (E) LAFAN1 mocap; 192 M steps
Lang-To-Loco ICLR 2026 OR k3Cyx3Uets G1 · 23 Diffusion student over MoE teacher 64-D motion latent (E); no vision MotionMillion (retargeting-free)
HWC-Loco ICLR 2026 OR 3UE3Aatcjy H1/G1 · 19 ZMP-constrained robust locomotion 19-DoF, VAE privileged-state CMU mocap; sim-only
UniFP CoRL 2025 ★ 2505.20829 B2-Z1/G1 · 18/29 Unified position+force loco-manip Force as first-class output Sim cmd combos; +39.5%
VideoMimic CoRL 2025 ★ 2505.03729 G1 · 23 Stairs/terrain/sit-stand via root cmds Drop target-angle conditioning Human video → 4D reconstruct → retarget
VIRAL CVPR 2026 2511.15200 G1 Zero-shot loco-manip Delta actions + RSI Teacher→RGB student, 64 GPUs; 54 cycles
Open-Sim-to-Real CVPR 2026 2512.01061 Humanoid Door-opening loco-manip Pure RGB→action Staged-reset teacher → GRPO; 31.7% > human
EgoVLA CVPR 2026 2507.12440 H1 + 2×12 hands Bimanual loco-manip Wrist SE(3) + MANO 500 K ego pairs; IK retarget; +43 pp long-horizon

★ = Best/Best-Student Paper. "OR" = OpenReview ID (no arXiv listed).


9. Decision guide

  1. Need true bipedal loco-manipulation (walk + manipulate)? → F1 latent-vocabulary VLA over an RL WBC (WholeBodyVLA, LeVERB). Accept that it's mostly sim-trained today.
  2. Stationary/upper-body humanoid, want best manipulation? → F2 dual-system (GR00T, Helix-style). Balance is not your problem; invest in the fast head and teleop.
  3. Balance/locomotion is the hard part, semantics secondary? → F3 controller substrate first (BFM-Zero, HWC-Loco, UniFP), then prompt it with a VLM.
  4. Contact-rich loco-manip (push doors, wipe, carry)? → unified position+force (UniFP) under the WBC.
  5. Data-constrained (no teleop farm)? → human-video / sim engine: VideoMimic or EgoScale-style ego-video for manipulation; VIRAL / Open-Sim-to-Real teacher-student for vision sim-to-real.
  6. Can't hit real-time with a big VLM? → async fast-slow (DuoCore-FS, Fast-in-Slow) and a high-rate System-1 head.

10. Limitations & open problems

  1. VLA-driven balance is still delegated, not learned end-to-end. Every credible loco-manip VLA puts an RL WBC underneath; nobody robustly emits balance-critical whole-body torques directly from a language-conditioned policy. Whether end-to-end is even desirable is open.
  2. The bipedal loco-manip VLA set is tiny and mostly simulated. LeVERB, WholeBodyVLA, HumanVLA — small sample, sim-heavy, few real-world long-horizon demos. Most "humanoid VLAs" are upper-body or mobile-base.
  3. Data is the binding constraint, and retargeting is fragile. Human-video transfer hinges on metric-scale reconstruction + IK under joint limits; EgoVLA shows human video alone gives 0% zero-shot. No standard humanoid-VLA benchmark exists (LeVERB-Bench, Isaac Humanoid Manip Benchmark, RoboCasa are not unified).
  4. Marketing ≠ capability. Figure Helix has no paper and is upper-body; vendor numbers are unverifiable. GR00T's later versions (N1.5+) are blog/model-card only, not peer-reviewed.
  5. High-DoF credit assignment is unsolved. HVD's value-decomposition helps offline RL but the broad problem — learning coordinated 30-DoF whole-body actions from limited data — remains hard.
  6. Compute walls. VIRAL shows vision sim-to-real for humanoids needs ~64 GPUs; this gates academic reproduction.
  7. Frontier, unverified (cite with care): 2026 preprints surfaced but not source-verified here — PhysiFlow (2603.05410, multi-brain latent flow-matching whole-body VLA), HEX (2604.07993, humanoid-aligned experts for cross-embodiment whole-body), HumanoidExo (2510.03022, exoskeleton-data whole-body VLA), Cybo-Waiter (2603.10675). Verify before relying on them.

11. Adjacent systems (not bipedal — included to prevent miscategorization)


12. Links


🗓 State of the Field (updated Aug 2026)

Verdict: the triple-system recipe (VLM + flow expert + RL lower body) became the open reference at RSS 2026; data efficiency, not data volume, is the winning argument — and evaluation lags arms by a generation.

📈 Trend

RSS 2026 was humanoid loco-manipulation's coming-out: Ψ₀ (open foundation model; 800 h video + 30 h robot data beats >10× co-trained corpora incl. GR00T N1.6 by >40 pp), HiWET (world-frame EE tracking beats body-frame), HAIC (dynamics-aware interaction WM), EgoHumanoid + HoMMI (robot-free demonstration pipelines), plus a deep whole-body-control bench (TeleGate, OmniXtreme, X-Loco, pixel locomotion).

⚖️ Approaches & trade-offs

Fork Options Trade-off
Data recipe Co-train human+humanoid in one policy (GR00T/EgoVLA/H-RDT line) vs decouple video→representation / robot→control (Ψ₀) Hardware evidence favors decoupling at 36-DoF scale — the humanoid face of the human-video fork
Collection interface Decoupled teleop rigs (PICO+MANUS+trackers; locomotion delegated) vs robot-free human interfaces (HoMMI's UMI+ego) Stability & fidelity vs scalability
Control stack Whole-body end-to-end vs triple-system (System-0 RL lower body) Expressiveness vs stability; RL controller caps agility (no dynamic bracing yet)

⚠️ Limitations & open problems

  • Per-task fine-tuning sits inside every published loop (Ψ₀: 80 demos/task) — no humanoid zero-shot generalist.
  • Payload and hand-DoF (Dex3-1-class) cap difficulty; precision insertion unsolved.
  • No humanoid OOD protocol or perturbation suite; intervention-assisted scoring is the norm.

Latest (preprint): ω-0 🆕 — whole-body latent-predictive WAM for concurrent loco-manipulation (81.8% on 11 household tasks vs 44.5% ψ-0). See Latest Papers.

← Back to Home

⚠️ **GitHub.com Fallback** ⚠️