ICLR 2026 WholeBodyVLA - Heungwoo/research GitHub Wiki
WholeBodyVLA — Unified Latent VLA for Whole-Body Loco-Manipulation (in-depth)
Venue: ICLR 2026 · arXiv: 2512.11047 · OpenReview: OCJmVjyzN7 · GitHub: OpenDriveLab/WholebodyVLA · Project: opendrivelab.com/wholebodyvla Authors: Haoran Jiang, Jin Chen, Qingwen Bu, Li Chen, Modi Shi, Yanjie Zhang, Delong Li, Chuanzhe Suo, Chuang Wang, Zhihui Peng, Hongyang Li (Fudan University · OpenDriveLab & MMLab, HKU · AgiBot · SII) Platform: AgiBot X2 bipedal humanoid (27 DoF) · Category: Humanoid whole-body VLA Companion: Humanoid VLA review
Figures below are embedded from the arXiv HTML (v1). If an image fails to load, the caption describes it.
TL;DR
WholeBodyVLA is one of the few true bipedal loco-manipulation VLAs — a policy that not only manipulates with two arms but also commands the legs to walk, turn, and squat so the robot can manipulate across a large workspace rather than from a fixed stance. Its three design moves answer the three humanoid-VLA pain points head-on:
- Data scarcity → a Latent Action Model (LAM) mines loco-manipulation priors from ~300 h of cheap action-free egocentric human video (a head-camera walkaround), not an expensive humanoid teleop farm.
- High-DoF whole-body action → the VLA emits two latent action codes (manipulation + locomotion) decoded by a lightweight head into upper-body joint angles + a discrete locomotion command — a unified latent interface, not 27 raw joint targets.
- Bipedal balance → a dedicated Loco-Manipulation-Oriented (LMO) RL controller turns the high-level locomotion command into stable advancing/turning/squatting, replacing the brittle velocity-tracking controllers that wobble under arm motion.
On AgiBot X2 it reaches 78.0% average success (mean over 6 subgoals), +21.3 points over the strongest VLA baseline (OpenVLA-OFT + LMO) and +36 points over GR00T + LMO. The standout ablation: removing the LAM pretraining costs ~38.7 points, and human-video pretraining gives the same success with ~8× fewer robot teleop trajectories.
1. The problem — why "large-space" loco-manipulation is hard

Figure 1 (teaser). The robot chains tasks that a fixed-base manipulator cannot: it walks up to a table, packs a bag, turns and steps to a shelf, loads a box, then pushes a cart. Every one of these requires the base to move in service of the manipulation goal.
A tabletop VLA assumes a fixed base — any arm motion is safe and the workspace is whatever the arms can reach. A humanoid breaks both assumptions:
- Manipulation-aware locomotion is missing. Prior humanoid stacks are either modular (a planner picks a spot, a separate locomotion controller walks there, then a manipulation policy runs in place) or end-to-end but trained only on in-place data. Neither learns to move the base as part of the manipulation, so the robot is confined to a small reachable region.
- Two root causes (the paper's framing): (1) humanoid teleoperation data is scarce — you cannot easily teleoperate a balancing biped at scale to acquire loco-manipulation knowledge; (2) existing RL locomotion controllers lack the precision/stability to faithfully execute the commands a manipulation policy would issue (they track velocities, drift, and destabilize when the arms swing).
WholeBodyVLA attacks both: cheap video for the data problem, a purpose-built RL controller for the execution problem, and a latent interface to glue them.
2. Method overview

Figure 2 (framework). Three components in one pipeline:
- LAM learns a discrete latent action vocabulary from raw video (no action labels).
- The VLA backbone reads the egocentric image + language and predicts the next manipulation latent and locomotion latent; a lightweight decoder grounds them to upper-body joint angles and a discrete locomotion command.
- The LMO RL policy consumes the locomotion command and produces stable whole-body leg/waist actions that keep the robot balanced while it manipulates.
2.1 Latent Action Model — priors from action-free video
The key data trick. A VQ-VAE with a DINOv2 feature backbone is trained on consecutive frame pairs (oₜ, oₜ₊ₖ): an encoder maps the transition to a continuous latent, quantizes it to a codebook entry cₜ, and a decoder reconstructs the future frame ôₜ₊ₖ = 𝒟(oₜ, cₜ) under an MSE reconstruction loss. The codebook entry becomes a pseudo-action — a compact token capturing "what changed between these two frames" — learnable from any video.
Crucial design choice — two separate LAMs, not one. Manipulation video comes from a (roughly) static head camera watching the hands; locomotion video comes from a moving camera as the person walks. These visual modalities conflict, so the paper trains a manipulation LAM (on the AgiBot World real-robot dataset) and a locomotion LAM (on ~300 h of self-collected head-camera walking video) separately. The ablation confirms this matters (see §4).
2.2 VLA backbone — predict two latent codes, decode to whole-body commands
The VLA (initialized from Prismatic-7B, the OpenVLA-family VLM) takes the egocentric image oₜ and language instruction ℓ and jointly predicts both latent codes: π_θ(cₜ^mani, cₜ^loco | oₜ, ℓ). A lightweight grounding decoder f then maps the predicted latents (plus robot state sₜ) to the actual command:
aₜ = f(ĉₜ^mani, ĉₜ^loco, sₜ) → (i) upper-body joint angles and (ii) a locomotion command for the LMO controller.
This is a latent-action-token VLA (discrete codes decoded to actions), not a flow-matching/diffusion head — and the dual-stream output (arms + locomotion) is the explicit whole-body split that most humanoid VLAs lack.
2.3 LMO RL policy — making the legs obey precisely
The locomotion command is discrete and intent-based, not a velocity target:
uₜ = [s_x, s_y, s_ψ, h*] ∈ {−1, 0, 1}³ × ℝ
— s_x / s_y / s_ψ are forward / lateral / yaw intents (back / stop / go), and h* is a target stance height (for squatting). Replacing velocity tracking with start–stop semantics is what gives precise "advance one step / turn in place / squat to this height" behavior.
- Observation (proprioception only, no privileged info):
Oₜ = [uₜ, ωₜ, gₜ, qₜ, q̇ₜ, aₜ₋₁]— command, base angular velocity, gravity vector, joint positions/velocities, previous action. - Reference shaping: a smooth tanh gate
v_k^ref(t) = v_k^goal · tanh[α(s_k − s̄_k(t))]with an exponentially-smoothed intent flag prevents impulsive accelerations (the cause of arm-induced instability). - Two-stage curriculum: Stage I learns a basic gait with randomized speeds; Stage II fixes cruising speeds and adds a directional-accuracy reward, structured arm-motion perturbations (so the legs stay stable while the arms work), and a stand-still penalty.
- Training: MuJoCo, 50 Hz control, single H100.
3. Data collection & hardware

Figure 4 (data pipeline). The locomotion prior comes from one operator with a head-mounted camera recording everyday advancing/turning/squatting — no robot, no teleop rig, no action labels. This is the cheapest possible data source and is the paper's main scalability claim.

Figure 5 (platform & teleop). AgiBot X2 is a 27-DoF bipedal humanoid: 2×7-DoF arms with Omnipicker grippers (14), 2×6-DoF legs (12), 1-DoF waist; head-mounted Intel RealSense D435i for egocentric RGB-D. Deployment splits compute — RTX 4090 runs the VLA, an onboard NanoPi runs the RL policy, linked over ZeroMQ.
Data summary:
- Pretraining (action-free): AgiBot World real-robot manipulation video (manipulation LAM) + ~300 h head-camera human video (locomotion LAM).
- Robot finetuning (teleop): only ~50 trajectories per task across 3 tasks (bag packing, box loading, cart pushing), collected via Meta Quest Pro VR + joystick.
- Training: LAMs 30k steps each (batch 256, 8×H100); VLA 20k pretrain + 10k finetune steps.
4. Results
Numbers verified against the arXiv HTML (v1) Table 2 & Table 3 — transcribed verbatim, averages independently re-derived. Earlier reconstructed per-task fractions have been replaced with the real per-subgoal values.
Main benchmark — Table 2 (3 tasks, each split into 2 subgoals; every subgoal scored over 25 trials; Avg = mean of the 6 subgoal success rates). Values transcribed verbatim from the arXiv HTML; averages re-derived and confirmed (e.g. WholeBodyVLA (23+13+19+17+23+22)/150 = 78.0%).
| Method | Bag: Grasp | Bag: Move&Squat | Box: Squat&Grasp | Box: Rise&Turn | Cart: Grab | Cart: Push | Avg |
|---|---|---|---|---|---|---|---|
| Modular Design | 22/25 | 12/25 | 9/25 | 9/25 | 22/25 | 22/25 | 64.0% |
| GR00T w/ LMO | 20/25 | 10/25 | 6/25 | 4/25 | 12/25 | 11/25 | 42.0% |
| OpenVLA-OFT w/ LMO | 19/25 | 6/25 | 12/25 | 12/25 | 22/25 | 14/25 | 56.7% |
| WholeBodyVLA (ours) | 23/25 | 13/25 | 19/25 | 17/25 | 23/25 | 22/25 | 78.0% |
Deltas: WholeBodyVLA beats OpenVLA-OFT + LMO by +21.3, GR00T + LMO by +36.0, and the modular planner-plus-controller pipeline by +14.0 — the unified latent approach wins over the traditional decomposition. The hardest subgoals are the manipulation-while-moving ones (Box "Squat&Grasp" / "Rise&Turn"): WholeBodyVLA holds 19/25 and 17/25 where GR00T collapses to 6/25 and 4/25.
Ablations (what each component is worth):
| Variant | Avg | Δ |
|---|---|---|
| Full WholeBodyVLA | 78.0% | — |
| w/ shared (single) LAM | 66.0% | −12.0 |
| w/ manipulation-only LAM | 63.3% | −14.7 |
| w/ velocity-based RL (instead of LMO) | 54.0% | −24.0 |
| w/o LAM (no pretraining) | 39.3% | −38.7 |
Reading: the LAM pretraining is the single biggest lever (−38.7 without it); the LMO controller beats a conventional velocity-RL controller by +24; and separating the two LAMs beats sharing one (+12) — confirming the manipulation/locomotion modality-conflict argument.
LMO ablation — Table 3 (MuJoCo). Locomotion accuracy is reported as Position / Quaternion error (mean ± std) for forward-backward, lateral, and turning; manipulation stability is CoM sway (CoMS) while standing / squatting (lower is better). Verbatim:
| Method | Fwd & Back (Pos./Quat.) | Left & Right (Pos./Quat.) | Turning (Pos./Quat.) | CoMS Stand | CoMS Squat |
|---|---|---|---|---|---|
| LMO (ours) | 0.21±0.01 / 0.05±0.01 | 0.55±0.01 / 0.06±0.01 | 0.05±0.01 / 0.19±0.01 | 0.03±0.02 | 0.03±0.02 |
| Vel.-based policy | 0.24±0.04 / 0.12±0.02 | 0.60±0.05 / 0.17±0.06 | 0.26±0.01 / 0.20±0.06 | 0.06±0.04 | 0.05±0.04 |
The LMO controller is more accurate on every axis — most strikingly turning position error 0.05 vs 0.26 and forward orientation 0.05 vs 0.12 — and halves CoM sway, with much tighter variance. (Note: the second number is a quaternion error, not radians.)

Figure 3 (scaling & generalization). The data-efficiency result is the most important: with 50% human-video pretraining, the model matches the no-pretrain variant using roughly 8× fewer robot teleop trajectories (~25 vs ~200). Extended-task panels show the approach generalizing to terrain traversal, long-horizon sequences, and visual navigation, consistently above the velocity-based baseline.
5. Significance — where it sits in the humanoid-VLA stack
WholeBodyVLA is, alongside LeVERB, the cleanest instance of the latent-vocabulary whole-body VLA family (Humanoid VLA review §3, Family F1): the VLA emits a learned latent and an RL whole-body controller owns balance. It is currently the strongest match to all three humanoid axes at once:
- Axis 1 (mobility/balance): the LMO controller is a genuine manipulation-aware locomotion contribution (advance/turn/squat with arm-perturbation training), not a generic walker.
- Axis 2 (high-DoF action): the dual latent code → dual-stream head is one of the few explicit arm-vs-locomotion action decompositions in the literature.
- Axis 3 (data): action-free egocentric video via LAM is the standout — the same "mine pseudo-actions from video" idea as GR00T's latent-action base, but pushed to locomotion priors from a head-camera walkaround.
The conceptual takeaway: you can get loco-manipulation knowledge without humanoid teleop if you (a) learn a latent action vocabulary from video and (b) hand execution to a controller built for precise start–stop locomotion.
6. Critical read & limitations
- Small real-task suite. The headline numbers come from 3 tasks with ~50 teleop trajectories each. Strong as a proof-of-concept; the breadth claims ("large-space", high extensibility) rest on the Figure-3 extended panels more than a large benchmark.
- The +21.3% is baseline-relative. It's over OpenVLA-OFT+LMO; absolute 78% on 3 tasks is good but not saturated, and the per-task spread (cart pushing easy, box loading hard) is wide.
- LMO trained/evaluated in MuJoCo at 50 Hz. Locomotion accuracy is reported in sim; real-robot balance is shown qualitatively (CoM sway) but there's no large real-world locomotion-precision study.
- Two-LAM design adds pipeline complexity and a modality-split assumption; the shared-LAM ablation (−12) shows it helps but also that a unified model is not hopeless.
- Latent grounding decoder is thin. Whether the lightweight
flimits fine manipulation dexterity (vs a diffusion/flow head) is untested here. - Reproducibility: code is released (OpenDriveLab/WholebodyVLA); the ~300 h human video and exact LAM checkpoints are the gating artifacts to watch.
7. Links & related
- arXiv: 2512.11047 · arXiv HTML (figures): html/2512.11047v1 · OpenReview: OCJmVjyzN7 · GitHub: OpenDriveLab/WholebodyVLA · Project: opendrivelab.com/wholebodyvla
- Companion reviews: Humanoid VLA (Family F1 context) · System 0/1/2 (the VLA→WBC handoff) · GR00T Series (latent-action-from-video precedent) · Cross-Embodiment
- Counterpart: UniHM (dexterous-hand whole-body)