ICLR 2026 WholeBodyVLA - Heungwoo/research GitHub Wiki

WholeBodyVLA — Unified Latent VLA for Whole-Body Loco-Manipulation (in-depth)

Venue: ICLR 2026 · arXiv: 2512.11047 · OpenReview: OCJmVjyzN7 · GitHub: OpenDriveLab/WholebodyVLA · Project: opendrivelab.com/wholebodyvla Authors: Haoran Jiang, Jin Chen, Qingwen Bu, Li Chen, Modi Shi, Yanjie Zhang, Delong Li, Chuanzhe Suo, Chuang Wang, Zhihui Peng, Hongyang Li (Fudan University · OpenDriveLab & MMLab, HKU · AgiBot · SII) Platform: AgiBot X2 bipedal humanoid (27 DoF) · Category: Humanoid whole-body VLA Companion: Humanoid VLA review

Figures below are embedded from the arXiv HTML (v1). If an image fails to load, the caption describes it.


TL;DR

WholeBodyVLA is one of the few true bipedal loco-manipulation VLAs — a policy that not only manipulates with two arms but also commands the legs to walk, turn, and squat so the robot can manipulate across a large workspace rather than from a fixed stance. Its three design moves answer the three humanoid-VLA pain points head-on:

  1. Data scarcity → a Latent Action Model (LAM) mines loco-manipulation priors from ~300 h of cheap action-free egocentric human video (a head-camera walkaround), not an expensive humanoid teleop farm.
  2. High-DoF whole-body action → the VLA emits two latent action codes (manipulation + locomotion) decoded by a lightweight head into upper-body joint angles + a discrete locomotion command — a unified latent interface, not 27 raw joint targets.
  3. Bipedal balance → a dedicated Loco-Manipulation-Oriented (LMO) RL controller turns the high-level locomotion command into stable advancing/turning/squatting, replacing the brittle velocity-tracking controllers that wobble under arm motion.

On AgiBot X2 it reaches 78.0% average success (mean over 6 subgoals), +21.3 points over the strongest VLA baseline (OpenVLA-OFT + LMO) and +36 points over GR00T + LMO. The standout ablation: removing the LAM pretraining costs ~38.7 points, and human-video pretraining gives the same success with ~8× fewer robot teleop trajectories.


1. The problem — why "large-space" loco-manipulation is hard

Figure 1 — teaser: AgiBot X2 performing consecutive loco-manipulation tasks (pack a bag, load a box, push a cart) by walking between stations.

Figure 1 (teaser). The robot chains tasks that a fixed-base manipulator cannot: it walks up to a table, packs a bag, turns and steps to a shelf, loads a box, then pushes a cart. Every one of these requires the base to move in service of the manipulation goal.

A tabletop VLA assumes a fixed base — any arm motion is safe and the workspace is whatever the arms can reach. A humanoid breaks both assumptions:

  • Manipulation-aware locomotion is missing. Prior humanoid stacks are either modular (a planner picks a spot, a separate locomotion controller walks there, then a manipulation policy runs in place) or end-to-end but trained only on in-place data. Neither learns to move the base as part of the manipulation, so the robot is confined to a small reachable region.
  • Two root causes (the paper's framing): (1) humanoid teleoperation data is scarce — you cannot easily teleoperate a balancing biped at scale to acquire loco-manipulation knowledge; (2) existing RL locomotion controllers lack the precision/stability to faithfully execute the commands a manipulation policy would issue (they track velocities, drift, and destabilize when the arms swing).

WholeBodyVLA attacks both: cheap video for the data problem, a purpose-built RL controller for the execution problem, and a latent interface to glue them.


2. Method overview

Figure 2 — framework: (left) dual Latent Action Model pretraining from action-free video, (middle) the VLA backbone predicting manipulation + locomotion latent codes, (right) the LMO RL low-level policy executing locomotion commands.

Figure 2 (framework). Three components in one pipeline:

  1. LAM learns a discrete latent action vocabulary from raw video (no action labels).
  2. The VLA backbone reads the egocentric image + language and predicts the next manipulation latent and locomotion latent; a lightweight decoder grounds them to upper-body joint angles and a discrete locomotion command.
  3. The LMO RL policy consumes the locomotion command and produces stable whole-body leg/waist actions that keep the robot balanced while it manipulates.

2.1 Latent Action Model — priors from action-free video

The key data trick. A VQ-VAE with a DINOv2 feature backbone is trained on consecutive frame pairs (oₜ, oₜ₊ₖ): an encoder maps the transition to a continuous latent, quantizes it to a codebook entry cₜ, and a decoder reconstructs the future frame ôₜ₊ₖ = 𝒟(oₜ, cₜ) under an MSE reconstruction loss. The codebook entry becomes a pseudo-action — a compact token capturing "what changed between these two frames" — learnable from any video.

Crucial design choice — two separate LAMs, not one. Manipulation video comes from a (roughly) static head camera watching the hands; locomotion video comes from a moving camera as the person walks. These visual modalities conflict, so the paper trains a manipulation LAM (on the AgiBot World real-robot dataset) and a locomotion LAM (on ~300 h of self-collected head-camera walking video) separately. The ablation confirms this matters (see §4).

2.2 VLA backbone — predict two latent codes, decode to whole-body commands

The VLA (initialized from Prismatic-7B, the OpenVLA-family VLM) takes the egocentric image oₜ and language instruction ℓ and jointly predicts both latent codes: π_θ(cₜ^mani, cₜ^loco | oₜ, ℓ). A lightweight grounding decoder f then maps the predicted latents (plus robot state sₜ) to the actual command:

aₜ = f(ĉₜ^mani, ĉₜ^loco, sₜ) → (i) upper-body joint angles and (ii) a locomotion command for the LMO controller.

This is a latent-action-token VLA (discrete codes decoded to actions), not a flow-matching/diffusion head — and the dual-stream output (arms + locomotion) is the explicit whole-body split that most humanoid VLAs lack.

2.3 LMO RL policy — making the legs obey precisely

The locomotion command is discrete and intent-based, not a velocity target:

uₜ = [s_x, s_y, s_ψ, h*] ∈ {−1, 0, 1}³ × ℝ

— s_x / s_y / s_ψ are forward / lateral / yaw intents (back / stop / go), and h* is a target stance height (for squatting). Replacing velocity tracking with start–stop semantics is what gives precise "advance one step / turn in place / squat to this height" behavior.

  • Observation (proprioception only, no privileged info): Oₜ = [uₜ, ωₜ, gₜ, qₜ, q̇ₜ, aₜ₋₁] — command, base angular velocity, gravity vector, joint positions/velocities, previous action.
  • Reference shaping: a smooth tanh gate v_k^ref(t) = v_k^goal · tanh[α(s_k − s̄_k(t))] with an exponentially-smoothed intent flag prevents impulsive accelerations (the cause of arm-induced instability).
  • Two-stage curriculum: Stage I learns a basic gait with randomized speeds; Stage II fixes cruising speeds and adds a directional-accuracy reward, structured arm-motion perturbations (so the legs stay stable while the arms work), and a stand-still penalty.
  • Training: MuJoCo, 50 Hz control, single H100.

3. Data collection & hardware

Figure 4 — low-cost egocentric data-collection pipeline: a single operator wearing a head-mounted camera walks/turns/squats to record ~300 h of action-free locomotion video.

Figure 4 (data pipeline). The locomotion prior comes from one operator with a head-mounted camera recording everyday advancing/turning/squatting — no robot, no teleop rig, no action labels. This is the cheapest possible data source and is the paper's main scalability claim.

Figure 5 — AgiBot X2 hardware and the Meta Quest Pro VR + joystick teleoperation setup used for the small robot-specific finetuning set.

Figure 5 (platform & teleop). AgiBot X2 is a 27-DoF bipedal humanoid: 2×7-DoF arms with Omnipicker grippers (14), 2×6-DoF legs (12), 1-DoF waist; head-mounted Intel RealSense D435i for egocentric RGB-D. Deployment splits compute — RTX 4090 runs the VLA, an onboard NanoPi runs the RL policy, linked over ZeroMQ.

Data summary:

  • Pretraining (action-free): AgiBot World real-robot manipulation video (manipulation LAM) + ~300 h head-camera human video (locomotion LAM).
  • Robot finetuning (teleop): only ~50 trajectories per task across 3 tasks (bag packing, box loading, cart pushing), collected via Meta Quest Pro VR + joystick.
  • Training: LAMs 30k steps each (batch 256, 8×H100); VLA 20k pretrain + 10k finetune steps.

4. Results

Numbers verified against the arXiv HTML (v1) Table 2 & Table 3 — transcribed verbatim, averages independently re-derived. Earlier reconstructed per-task fractions have been replaced with the real per-subgoal values.

Main benchmark — Table 2 (3 tasks, each split into 2 subgoals; every subgoal scored over 25 trials; Avg = mean of the 6 subgoal success rates). Values transcribed verbatim from the arXiv HTML; averages re-derived and confirmed (e.g. WholeBodyVLA (23+13+19+17+23+22)/150 = 78.0%).

Method Bag: Grasp Bag: Move&Squat Box: Squat&Grasp Box: Rise&Turn Cart: Grab Cart: Push Avg
Modular Design 22/25 12/25 9/25 9/25 22/25 22/25 64.0%
GR00T w/ LMO 20/25 10/25 6/25 4/25 12/25 11/25 42.0%
OpenVLA-OFT w/ LMO 19/25 6/25 12/25 12/25 22/25 14/25 56.7%
WholeBodyVLA (ours) 23/25 13/25 19/25 17/25 23/25 22/25 78.0%

Deltas: WholeBodyVLA beats OpenVLA-OFT + LMO by +21.3, GR00T + LMO by +36.0, and the modular planner-plus-controller pipeline by +14.0 — the unified latent approach wins over the traditional decomposition. The hardest subgoals are the manipulation-while-moving ones (Box "Squat&Grasp" / "Rise&Turn"): WholeBodyVLA holds 19/25 and 17/25 where GR00T collapses to 6/25 and 4/25.

Ablations (what each component is worth):

Variant Avg Δ
Full WholeBodyVLA 78.0% —
w/ shared (single) LAM 66.0% −12.0
w/ manipulation-only LAM 63.3% −14.7
w/ velocity-based RL (instead of LMO) 54.0% −24.0
w/o LAM (no pretraining) 39.3% −38.7

Reading: the LAM pretraining is the single biggest lever (−38.7 without it); the LMO controller beats a conventional velocity-RL controller by +24; and separating the two LAMs beats sharing one (+12) — confirming the manipulation/locomotion modality-conflict argument.

LMO ablation — Table 3 (MuJoCo). Locomotion accuracy is reported as Position / Quaternion error (mean ± std) for forward-backward, lateral, and turning; manipulation stability is CoM sway (CoMS) while standing / squatting (lower is better). Verbatim:

Method Fwd & Back (Pos./Quat.) Left & Right (Pos./Quat.) Turning (Pos./Quat.) CoMS Stand CoMS Squat
LMO (ours) 0.21±0.01 / 0.05±0.01 0.55±0.01 / 0.06±0.01 0.05±0.01 / 0.19±0.01 0.03±0.02 0.03±0.02
Vel.-based policy 0.24±0.04 / 0.12±0.02 0.60±0.05 / 0.17±0.06 0.26±0.01 / 0.20±0.06 0.06±0.04 0.05±0.04

The LMO controller is more accurate on every axis — most strikingly turning position error 0.05 vs 0.26 and forward orientation 0.05 vs 0.12 — and halves CoM sway, with much tighter variance. (Note: the second number is a quaternion error, not radians.)

Figure 3 — generalization & data scaling: (a–b) success vs. number of teleop trajectories with/without human-video pretraining; (c) extended tasks (terrain traversal, long-horizon sequences, visual navigation).

Figure 3 (scaling & generalization). The data-efficiency result is the most important: with 50% human-video pretraining, the model matches the no-pretrain variant using roughly 8× fewer robot teleop trajectories (~25 vs ~200). Extended-task panels show the approach generalizing to terrain traversal, long-horizon sequences, and visual navigation, consistently above the velocity-based baseline.


5. Significance — where it sits in the humanoid-VLA stack

WholeBodyVLA is, alongside LeVERB, the cleanest instance of the latent-vocabulary whole-body VLA family (Humanoid VLA review §3, Family F1): the VLA emits a learned latent and an RL whole-body controller owns balance. It is currently the strongest match to all three humanoid axes at once:

  • Axis 1 (mobility/balance): the LMO controller is a genuine manipulation-aware locomotion contribution (advance/turn/squat with arm-perturbation training), not a generic walker.
  • Axis 2 (high-DoF action): the dual latent code → dual-stream head is one of the few explicit arm-vs-locomotion action decompositions in the literature.
  • Axis 3 (data): action-free egocentric video via LAM is the standout — the same "mine pseudo-actions from video" idea as GR00T's latent-action base, but pushed to locomotion priors from a head-camera walkaround.

The conceptual takeaway: you can get loco-manipulation knowledge without humanoid teleop if you (a) learn a latent action vocabulary from video and (b) hand execution to a controller built for precise start–stop locomotion.


6. Critical read & limitations

  • Small real-task suite. The headline numbers come from 3 tasks with ~50 teleop trajectories each. Strong as a proof-of-concept; the breadth claims ("large-space", high extensibility) rest on the Figure-3 extended panels more than a large benchmark.
  • The +21.3% is baseline-relative. It's over OpenVLA-OFT+LMO; absolute 78% on 3 tasks is good but not saturated, and the per-task spread (cart pushing easy, box loading hard) is wide.
  • LMO trained/evaluated in MuJoCo at 50 Hz. Locomotion accuracy is reported in sim; real-robot balance is shown qualitatively (CoM sway) but there's no large real-world locomotion-precision study.
  • Two-LAM design adds pipeline complexity and a modality-split assumption; the shared-LAM ablation (−12) shows it helps but also that a unified model is not hopeless.
  • Latent grounding decoder is thin. Whether the lightweight f limits fine manipulation dexterity (vs a diffusion/flow head) is untested here.
  • Reproducibility: code is released (OpenDriveLab/WholebodyVLA); the ~300 h human video and exact LAM checkpoints are the gating artifacts to watch.

7. Links & related

← Back to ICLR-2026 · Home