Review Humanoid VLA - Heungwoo/research GitHub Wiki
Compiled June 2026 · Focus: what makes a Vision-Language-Action model for a humanoid technically different from a tabletop-arm VLA — specifically (1) mobility / bipedal balance during manipulation, (2) the high-dimensional whole-body action space, and (3) data collection when you cannot simply teleoperate a balancing biped.
Companion reviews: System 0/1/2 architectures · GR00T series · Dexterous Manipulation · Cross-Embodiment · VLA Architectures · DuoCore-FS.
A tabletop VLA maps (image, instruction) → end-effector action. A humanoid cannot be modeled that way, for three reasons that organize this whole review:
- The robot can fall over. Every action is conditioned on keeping a ~19–29-DoF underactuated biped balanced. Almost no system lets the VLA emit balance-critical torques directly — balance is delegated to a separate RL whole-body controller (WBC) and the VLA talks to it through a latent or a command.
- The action space is 2–5× larger and heterogeneous. Arms + hands + waist + neck + legs, mixing position, force, and locomotion commands at different rates. The field's answers are latent action vocabularies, dual-system VLM→fast-policy splits, and stream-specific tokenization — not one monolithic head.
- You cannot cheaply teleoperate a walking humanoid. So the data engine shifts off teleop toward human video → reconstruction → retargeting, motion-capture → RL-tracking → distillation, and sim-RL teacher → vision student. Teleop survives mostly for stationary upper-body work (Helix, GR00T's real layer).
The single most important structural fact: "humanoid VLA" today is a stack, cleanly described by the System 0/1/2 framing — a slow VLM (System 2, 1–10 Hz), a fast visuomotor policy (System 1, 50–200 Hz), and a reflex/balance layer (System 0, 500 Hz–1 kHz). The genuinely-new research is in how the VLA hands off to the balance layer and where the training data comes from.
Headline caveat — read before believing any "whole-body" claim: several flagship "humanoid VLAs" are upper-body only and never walk — Figure Helix (35-DoF, upper body, no paper, vendor numbers only) and GR00T's public GR-1 deployment (stationary bimanual). True bipedal loco-manipulation VLAs — where the policy commands walking/balance — remain rare and mostly simulation-trained: LeVERB, WholeBodyVLA, HumanVLA, Humanoid-VLA.
| Tabletop-VLA assumption | Why a humanoid violates it | Field's response |
|---|---|---|
| The base is fixed; actions don't threaten stability | Underactuated biped; any arm motion shifts CoM and can topple it | Delegate balance to an RL WBC (System 0/1); VLA emits latent/command, not torque |
| ≤7–14 DoF, one action modality (EE pose) | 19–36 DoF spanning legs/waist/arms/hands/neck; mixes position + force + gait | Latent action vocab · dual-system split · stream-specific tokenization · hierarchical value decomposition |
| Teleop scales data linearly | A walking biped is unsafe/awkward to teleoperate; balance data isn't demonstrable | Human-video reconstruction · mocap→RL→distill · sim teacher→vision student |
| One control rate (~10–50 Hz) | Reasoning (1–10 Hz) and balance (≥500 Hz) differ by ~100× | Asynchronous fast-slow inference; per-rate heads |
| One embodiment per policy | Many humanoid morphologies (G1, H1, GR-1, Atlas, AgiBot, Astribot) | Per-embodiment MLP projection to a shared latent; cross-embodiment pyramids |
This is why humanoid VLA looks like "a VLM bolted onto a learned whole-body controller" rather than "scale imitation until it works."
The three-layer stack — a slow VLM hands off to a fast policy, which hands off to a balance layer:
flowchart TB
VLM["System 2 · slow VLM · 1–10 Hz<br/>latent / subtask / motion code"]
POL["System 1 · fast policy · 50–200 Hz<br/>diffusion / flow / AR head"]
WBC["System 0 · reflex & balance · 500 Hz–1 kHz<br/>RL whole-body controller"]
ROB["humanoid joints · 19–36 DoF"]
VLM --> POL --> WBC --> ROB
The four families — which approach fits your constraint:
flowchart TB
Q{Which family?}
Q -- learned-latent VLA over an RL WBC --> F1[F1. Latent-vocabulary whole-body VLA<br/>LeVERB · WholeBodyVLA]
Q -- VLM then fast-policy split --> F2[F2. Dual-system manipulation VLA<br/>GR00T · Helix · DuoCore-FS · Galaxea]
Q -- language to motion / limited vision --> F3[F3. Language-to-motion controllers<br/>Lang-To-Loco · BFM-Zero · Humanoid-GPT]
Q -- sim or human-video data engine --> F4[F4. Data & sim-to-real engines<br/>VideoMimic · VIRAL · EgoScale · Open-Sim-to-Real]
- F1 — Latent-vocabulary whole-body VLA. The VLA emits a learned latent action token; an RL WBC decodes it into balanced whole-body joint targets. The cleanest answer to all three axes at once. (LeVERB, WholeBodyVLA.)
- F2 — Dual-system manipulation VLA. Slow VLM + fast policy, heterogeneous DoF handled by per-embodiment MLPs. Strong on manipulation; balance is usually out of scope (upper-body or mobile-base). (GR00T N1–N1.7, Helix, DuoCore-FS, Galaxea G0, Fast-in-Slow.)
- F3 — Language-to-motion / promptable controllers. Bridge language (sometimes not vision) to whole-body motion via latents; these are often the System 0/1 substrate a full VLA sits on. (Lang-To-Loco/RoboGhost, BFM-Zero, Humanoid-GPT, HWC-Loco, HVD, LIFT, UniFP.)
- F4 — Data & sim-to-real engines. Not policies per se but the data-collection machinery that makes the others trainable. (VideoMimic, VIRAL, Open-Sim-to-Real, Gallant, EgoVLA, EgoScale.)
The defining humanoid problem. Three sub-patterns dominate:
The VLA almost never outputs balance torques. Instead a low-level RL WBC owns stability and the VLA steers it:
- LeVERB (arXiv 2506.13751, Berkeley): a System-1 RL WBC produces dynamics-feasible walking/turning/sitting from a latent instruction; the System-2 vision-language policy never touches joint torques.
- WholeBodyVLA (ICLR 2026, 2512.11047, AgiBot X2): a Loco-Manipulation-Oriented (LMO) RL policy for advancing / turning / squatting — "manipulation-aware locomotion" that moves the base to serve the manipulation goal. This is the explicit bipedal-loco-manip contribution.
- HumanVLA (NeurIPS 2024, 2406.19972): a state-based teacher trained with goal-conditioned RL + Adversarial Motion Prior (AMP) learns balanced walk-and-carry, then distilled into a VLA student.
- BFM-Zero (ICLR 2026, 2511.04131, Unitree G1, 29-DoF): a promptable Forward-Backward foundation controller with disturbance rejection (recovers from kicks/pushes never trained on) — the substrate a VLA prompts.
These are not VLAs but the balance layer the field depends on:
- HWC-Loco (ICLR 2026, H1/G1, 19-DoF): 3-stage hierarchy with a ZMP-feasibility constraint and robust optimization over an uncertainty set; survives ≤200 N pushes, 15 cm stairs, 20° slopes. Learned ZMP-threshold switching beats fixed thresholds (195 vs 459 switches).
- UniFP (CoRL 2025 Best Paper, 2505.20829): a single unified policy for position and force in legged loco-manipulation — infers external force from proprioception (no F/T sensor), no hand-engineered mode switching; +39.5% on contact-rich tasks (wiping, drawers).
- Gallant (CVPR 2026, 2511.14625): 3-D voxel-grid terrain perception (overhead + lateral constraints, not just a 2-D heightmap) for navigating pipes/shelves/stairs; near-100% staircase success.
Reasoning, control, and balance run ~100× apart: VLM 1–10 Hz → fast policy 50–200 Hz → balance 500 Hz–1 kHz. Industrial systems make this explicit — Figure's reported System-0 is a 1 kHz neural prior trained on >1,000 h of human motion capture that replaced ~109k lines of C++ balance code (per Review-System-0-1-2). The async-inference papers (DuoCore-FS 2.6×, Fast-in-Slow 117.7 Hz) exist precisely to keep the fast layer fed while the slow VLM lags.
Reality check: the most-publicized "humanoid" VLAs sidestep Axis 1 entirely. Helix is upper-body only (no walking/balance claim). DuoCore-FS explicitly omits a System-0 balance layer (quasi-static kiosk). GR00T N1's GR-1 deployment is stationary. Bipedal balance during VLA-driven manipulation is demonstrated by a short list — LeVERB, WholeBodyVLA, HumanVLA — and mostly in simulation.
DoF roughly doubles-to-quintuples vs a 7-DoF arm: H1 ~19, G1 ~23–29, Astribot S1 25, Figure 35 (upper only), Dexora 36 (bimanual+hands). Five distinct strategies:
| Strategy | Mechanism | Representative | Why it helps |
|---|---|---|---|
| A. Latent action vocabulary | VLA emits a learned latent token; RL WBC decodes to whole-body joints | LeVERB, WholeBodyVLA | Decouples semantics from the 30-DoF control problem; no hand-crafted action primitives |
| B. Dual-system + per-embodiment projection | Slow VLM → shared latent → per-embodiment MLP encoders/decoders → fast head | GR00T N1–N1.7, Helix, Galaxea G0, Fast-in-Slow | One model spans many morphologies; no explicit upper/lower split needed |
| C. Stream-specific tokenization | Residual-VQ-VAE with parallel codebooks per stream (position / SO(3) / gripper) | DuoCore-FS (29-dim → 36 fixed tokens, 3.4× more compact than FAST) | Fixed-length tokens make AR decoding of whole-body actions tractable and fast |
| D. Hierarchical value decomposition | Decompose the value function along kinematic structure (not the policy) | HVD (WB-50 dataset) | Credit assignment in high-DoF offline RL while keeping one unified policy |
| E. Behavioral latent / FB representation | Map 29-DoF control into a low-dim latent z (e.g. R²⁵⁶); plan/track in latent |
BFM-Zero, Lang-To-Loco (64-D motion latent) | Few-shot CEM/MPC and language-conditioning become low-dim problems |
Notes that matter:
- Explicit upper/lower-body decomposition is rare. Most systems use a shared latent and let projection layers sort out morphology. WholeBodyVLA's dual-stream head (separate arm-joint vs locomotion-command outputs) is one of the few explicit splits.
- Frequency mismatch is part of the action-space problem. Helix runs S1 at 200 Hz (80M params) under a 7–9 Hz 7B VLM; FiS reuses only the last 2 of 32 LLaMA blocks at high rate; DuoCore-FS runs an async 25–30 Hz fast head under a 1–3 Hz slow head.
- Dexterity rides on top. Bimanual 12-DoF hands (EgoVLA's Inspire hands, Dexora's XHAND) compound the DoF count — see Dexterous Manipulation for the hand-specific story; the whole-body papers mostly treat hands as another projected stream.
You cannot teleoperate a balancing biped at scale, so humanoid VLA data comes from four routes. Retargeting (human morphology → robot URDF under joint limits) is the recurring crux.
The dominant academic route — cheap monocular/egocentric human video lifted to robot actions:
- VideoMimic (CoRL 2025 Best Student Paper, 2505.03729, G1): everyday RGB video → 4-D reconstruction (SMPL mesh + scene, metric scale via joint Levenberg-Marquardt optimization) → retarget → 4-stage sim (mocap pretrain → scene-conditioned tracking → DAgger distill → under-conditioned RL). Removing the mocap-pretrain stage breaks it.
- EgoVLA (CVPR 2026, 2507.12440, H1 + 2×12-DoF Inspire): ~500K egocentric image-action pairs (HOI4D/HOT3D/HoloAssist/TACO), action = wrist SE(3) + MANO hand params, deployed via IK + hand retargeting. Pretraining lifts long-horizon success 2.22% → 45.93% vs ACT. Zero-shot (no robot finetune) = 0% — human video alone is insufficient.
-
EgoScale (arXiv 2602.16710): 20,854 h of action-labeled egocentric video (monocular SLAM + 21-keypoint hand pose + CasADi/IPOPT retargeting) — ~20× prior ego-video efforts; feeds GR00T N1.7 with a log-linear scaling law
L_val = 0.024 − 0.003·ln(D)(R²=0.998). - WholeBodyVLA learns a Latent Action Model from action-free egocentric video — no demonstration labels — then a thin teleop layer grounds it.
- Humanoid-VLA (arXiv 2502.14795): language-motion pre-alignment → egocentric video-conditioned PEFT → self-supervised pseudo-annotation of unlabeled video.
For locomotion/whole-body motion (where balance, not semantics, is the target):
- Humanoid-GPT (CVPR 2026, G1): 2-billion-frame retargeted mocap (AMASS + LAFAN1 + Motion-X++ + MotionMillion + in-house), Harmonic-Motion-Embedding diversity clustering → per-cluster RL experts → DAgger into one causal GPT; zero-shot motion tracking at <1.5 ms (TensorRT).
- BFM-Zero: LAFAN1 (40 motions retargeted to G1), EMD-prioritized sampling, 192 M env-steps off-policy.
- Lang-To-Loco (RoboGhost): MotionMillion (50,378 seqs → 3,261 stable after filtering); language → 64-D motion latent → diffusion student; retargeting-free.
Privileged-state RL in sim, distilled to an RGB policy that transfers zero-shot:
- VIRAL (CVPR 2026, 2511.15200, NVIDIA/CMU/Berkeley, G1): privileged teacher → RGB student via large-scale tiled-rendering DAgger; compute is decisive (up to 64 GPUs; low-compute runs fail); 54 zero-shot real cycles.
- Open-Sim-to-Real (CVPR 2026, 2512.01061): staged-reset PPO teacher → RGB student → GRPO finetune for partial observability; humanoid door-opening 31.7% faster than human teleoperators.
- LeVERB and HumanVLA are entirely synthetic (rendered kinematic demos + RL); no teleop, no real video.
- Figure Helix: ~500 h multi-robot multi-operator teleop (vendor figure; no paper).
- GR00T: the "data pyramid" — 88 h GR-1 teleop (VIVE + Xsens) at N1, scaling to thousands of hours of real teleop by N1.6 (YAM/AGIBot/G1). The version-over-version story is which layer gets scaled: N1→N1.5 synthetic (DreamGen) · N1.5→N1.6 real-teleop · N1.6→N1.7 human-video (EgoScale).
- Dexora: hybrid exoskeleton backpack + Apple Vision Pro teleop (gross arm + markerless fingers); 100 K sim + 10 K real episodes.
- DuoCore-FS: just 10.22 h teleop, single kiosk task — efficient but narrow.
The data takeaway: academic whole-body VLAs lean synthetic / human-video; industrial manipulation VLAs lean teleop; GR00T is the only line that fuses all three at scale. Retargeting quality (metric-scale reconstruction, IK under joint limits) is the silent determinant of whether human-video data transfers.
- CoRL 2025 — human-video-as-cross-embodiment-source goes mainstream. VideoMimic (Best Student Paper) and UniFP (Best Paper) frame the year: monocular human video → balanced humanoid skills, and unified position-force loco-manipulation.
- NeurIPS 2025 — async dual-system inference. Fast-in-Slow embeds the fast policy inside the VLM's last blocks (117.7 Hz) — the architectural enabler for running a slow reasoner over a fast humanoid.
- ICLR 2026 — the whole-body controller substrate matures. A dense cluster of non-VLA controllers a VLA sits on: BFM-Zero (promptable FB foundation model), WholeBodyVLA (the standout true loco-manip VLA), Lang-To-Loco, HWC-Loco, HVD, LIFT.
- CVPR 2026 — the sim-to-real data engine. VIRAL, Open-Sim-to-Real, Gallant, Humanoid-GPT, EgoVLA — the vision venue owns how you generate humanoid training signal at scale (teacher-student, voxel terrain, mocap pretraining, ego-video).
- Industry (no papers) — dual-system + teleop. Figure Helix and NVIDIA GR00T define the deployed template; Helix is upper-body, GR00T is cross-embodiment with a real-teleop core.
- IROS 2026 🆕 — bimanual coordination structure + whole-body unification. The frontier is coordination architecture, not a second arm: 3D FlowMatch Actor (CMU/NVIDIA — one 3D policy for single and dual-arm, +41.4% PerAct2, beats 1000×-larger models) and EquiBim (bilateral symmetry-equivariance as an inductive bias). Whole-body: ULTRA (unified multimodal whole-body loco-manip), CEER (compliant EE+root unified interface), OmniDP (beyond-FOV omnidirectional 3D perception), DreamMimic (loco-manip via a world model), and a MoE-VLA for humanoid loco-manip (Responsibility-Induced Specialized Experts — the multi-task interference fix reaching whole-body). Data is the bottleneck → RL-based bimanual data generation + physics-informed retargeting (SPIDER). Context: IROS 2026 survey §5.3.
| System | Venue | arXiv/ID | Platform · DoF | Axis 1 — Mobility/Balance | Axis 2 — Action space | Axis 3 — Data |
|---|---|---|---|---|---|---|
| LeVERB | arXiv Jun 2025 | 2506.13751 | G1 · ~23–29 | RL WBC decodes latent → walk/turn/sit | Latent action vocabulary (A) | Synthetic only; LeVERB-Bench 58.5% |
| WholeBodyVLA | ICLR 2026 | 2512.11047 | AgiBot X2 · bipedal | LMO RL policy (advance/turn/squat) | Unified latent + dual-stream head (A) | Action-free ego video + thin teleop; +21.3% |
| HumanVLA | NeurIPS 2024 | 2406.19972 | Sim humanoid | AMP teacher → balanced walk-carry | Single whole-body policy | Sim teacher→VLA student; Human-in-the-Room set |
| Humanoid-VLA | arXiv Feb 2025 | 2502.14795 | Humanoid (unspec) | Whole-body control backbone | — | Language-motion align + ego video + pseudo-labels |
| GR00T N1–N1.7 | NVIDIA | 2503.14734 (N1) | GR-1/G1/Atlas… | N1 manipulation-only; G1 WBC in later code | Dual-system DiT + per-embodiment MLP (B) | Data pyramid (teleop+sim+ego); N1.7 = 20,854 h EgoScale |
| Figure Helix | blog only | no paper | Figure 02 · 35 (upper) | None (upper-body only) | Dual-system 7B@7–9 Hz + 80M@200 Hz (B) | ~500 h teleop (vendor) |
| DuoCore-FS | arXiv Dec 2025 | 2512.20188 | Astribot S1 · 25 | Omitted (quasi-static) | RVQ stream tokenizer (C); async 2.6× | 10.22 h teleop, 1 task |
| Galaxea G0 | arXiv Sep 2025 | 2509.00576 | Mobile base | Mobile (wheeled, not bipedal) | Dual-system planner+executor (B) | Open dataset; 3-stage curriculum |
| BFM-Zero | ICLR 2026 | 2511.04131 | G1 · 29 | Promptable FB controller; push recovery | FB latent z∈R²⁵⁶ (E) | LAFAN1 mocap; 192 M steps |
| Lang-To-Loco | ICLR 2026 | OR k3Cyx3Uets | G1 · 23 | Diffusion student over MoE teacher | 64-D motion latent (E); no vision | MotionMillion (retargeting-free) |
| HWC-Loco | ICLR 2026 | OR 3UE3Aatcjy | H1/G1 · 19 | ZMP-constrained robust locomotion | 19-DoF, VAE privileged-state | CMU mocap; sim-only |
| UniFP | CoRL 2025 ★ | 2505.20829 | B2-Z1/G1 · 18/29 | Unified position+force loco-manip | Force as first-class output | Sim cmd combos; +39.5% |
| VideoMimic | CoRL 2025 ★ | 2505.03729 | G1 · 23 | Stairs/terrain/sit-stand via root cmds | Drop target-angle conditioning | Human video → 4D reconstruct → retarget |
| VIRAL | CVPR 2026 | 2511.15200 | G1 | Zero-shot loco-manip | Delta actions + RSI | Teacher→RGB student, 64 GPUs; 54 cycles |
| Open-Sim-to-Real | CVPR 2026 | 2512.01061 | Humanoid | Door-opening loco-manip | Pure RGB→action | Staged-reset teacher → GRPO; 31.7% > human |
| EgoVLA | CVPR 2026 | 2507.12440 | H1 + 2×12 hands | Bimanual loco-manip | Wrist SE(3) + MANO | 500 K ego pairs; IK retarget; +43 pp long-horizon |
★ = Best/Best-Student Paper. "OR" = OpenReview ID (no arXiv listed).
- Need true bipedal loco-manipulation (walk + manipulate)? → F1 latent-vocabulary VLA over an RL WBC (WholeBodyVLA, LeVERB). Accept that it's mostly sim-trained today.
- Stationary/upper-body humanoid, want best manipulation? → F2 dual-system (GR00T, Helix-style). Balance is not your problem; invest in the fast head and teleop.
- Balance/locomotion is the hard part, semantics secondary? → F3 controller substrate first (BFM-Zero, HWC-Loco, UniFP), then prompt it with a VLM.
- Contact-rich loco-manip (push doors, wipe, carry)? → unified position+force (UniFP) under the WBC.
- Data-constrained (no teleop farm)? → human-video / sim engine: VideoMimic or EgoScale-style ego-video for manipulation; VIRAL / Open-Sim-to-Real teacher-student for vision sim-to-real.
- Can't hit real-time with a big VLM? → async fast-slow (DuoCore-FS, Fast-in-Slow) and a high-rate System-1 head.
- VLA-driven balance is still delegated, not learned end-to-end. Every credible loco-manip VLA puts an RL WBC underneath; nobody robustly emits balance-critical whole-body torques directly from a language-conditioned policy. Whether end-to-end is even desirable is open.
- The bipedal loco-manip VLA set is tiny and mostly simulated. LeVERB, WholeBodyVLA, HumanVLA — small sample, sim-heavy, few real-world long-horizon demos. Most "humanoid VLAs" are upper-body or mobile-base.
- Data is the binding constraint, and retargeting is fragile. Human-video transfer hinges on metric-scale reconstruction + IK under joint limits; EgoVLA shows human video alone gives 0% zero-shot. No standard humanoid-VLA benchmark exists (LeVERB-Bench, Isaac Humanoid Manip Benchmark, RoboCasa are not unified).
- Marketing ≠ capability. Figure Helix has no paper and is upper-body; vendor numbers are unverifiable. GR00T's later versions (N1.5+) are blog/model-card only, not peer-reviewed.
- High-DoF credit assignment is unsolved. HVD's value-decomposition helps offline RL but the broad problem — learning coordinated 30-DoF whole-body actions from limited data — remains hard.
- Compute walls. VIRAL shows vision sim-to-real for humanoids needs ~64 GPUs; this gates academic reproduction.
- Frontier, unverified (cite with care): 2026 preprints surfaced but not source-verified here — PhysiFlow (2603.05410, multi-brain latent flow-matching whole-body VLA), HEX (2604.07993, humanoid-aligned experts for cross-embodiment whole-body), HumanoidExo (2510.03022, exoskeleton-data whole-body VLA), Cybo-Waiter (2603.10675). Verify before relying on them.
- ByteDance GR-2 / GR-3 (2410.06158 / 2507.15493) — strong generative VLAs but bimanual tabletop, not bipedal.
- Hume (2505.21432) — System-2 value-query VLA; general dexterous manipulation, not humanoid-specific.
- Dexora (2605.18722) — 36-DoF bimanual dexterous VLA; fixed-base, no locomotion (see Dexterous Manipulation).
- Companion reviews: System 0/1/2 · GR00T Series · DuoCore-FS · Fast-in-Slow · Dexterous Manipulation · Cross-Embodiment · VLA Architectures
- Whole-body VLAs: WholeBodyVLA · BFM-Zero · Lang-To-Loco
- Controllers / substrate: HWC-Loco · HVD · LIFT · UniFP
- Data / sim-to-real engines: VideoMimic · VIRAL · Open-Sim-to-Real · Gallant · EgoVLA
Verdict: the triple-system recipe (VLM + flow expert + RL lower body) became the open reference at RSS 2026; data efficiency, not data volume, is the winning argument — and evaluation lags arms by a generation.
RSS 2026 was humanoid loco-manipulation's coming-out: Ψ₀ (open foundation model; 800 h video + 30 h robot data beats >10× co-trained corpora incl. GR00T N1.6 by >40 pp), HiWET (world-frame EE tracking beats body-frame), HAIC (dynamics-aware interaction WM), EgoHumanoid + HoMMI (robot-free demonstration pipelines), plus a deep whole-body-control bench (TeleGate, OmniXtreme, X-Loco, pixel locomotion).
| Fork | Options | Trade-off |
|---|---|---|
| Data recipe | Co-train human+humanoid in one policy (GR00T/EgoVLA/H-RDT line) vs decouple video→representation / robot→control (Ψ₀) | Hardware evidence favors decoupling at 36-DoF scale — the humanoid face of the human-video fork |
| Collection interface | Decoupled teleop rigs (PICO+MANUS+trackers; locomotion delegated) vs robot-free human interfaces (HoMMI's UMI+ego) | Stability & fidelity vs scalability |
| Control stack | Whole-body end-to-end vs triple-system (System-0 RL lower body) | Expressiveness vs stability; RL controller caps agility (no dynamic bracing yet) |
- Per-task fine-tuning sits inside every published loop (Ψ₀: 80 demos/task) — no humanoid zero-shot generalist.
- Payload and hand-DoF (Dex3-1-class) cap difficulty; precision insertion unsolved.
- No humanoid OOD protocol or perturbation suite; intervention-assisted scoring is the norm.
Latest (preprint): ω-0 🆕 — whole-body latent-predictive WAM for concurrent loco-manipulation (81.8% on 11 household tasks vs 44.5% ψ-0). See Latest Papers.
← Back to Home