Review Cross Embodiment - Heungwoo/research GitHub Wiki
Compiled April 2026 Β· Focus: how recent VLA papers (Ο-series + ICLR 2026 + CoRL 2025 + 2024β2026 arXiv) train one policy for many bodies, and how those designs compare.
This is a cross-paper review page organized by how the training data is mixed. For the deployment cut β which papers run one frozen checkpoint on multiple robots at inference (routed seen-robot vs zero-shot-to-unseen) β see the companion Single-Checkpoint Multi-Robot Deployment. For per-paper summaries see the linked Cross-embodiment pages. Related reviews: Review-VLA-Memory, Review-pi07, Review-VLM4VLA.
Every robot lab collects data on its own hardware. Every VLA needs to run on the hardware the customer has. The gap between those two is the cross-embodiment problem:
- Arm / gripper differences β different reach, DoF, force envelope, end-effector geometry.
- Base differences β fixed vs. wheeled vs. legged vs. humanoid, changes the whole workspace.
- Sensor differences β wrist cam vs. head cam; 3 views vs. 4.
- Control differences β joint space vs. end-effector space; 20 Hz vs. 50 Hz.
- Strategy differences β a bigger / heavier arm may need a completely different grasp strategy for the same task.
The headline 2026 result: Ο0.7 zero-shots laundry folding on a UR5e bimanual that was never in its laundry training data, matching expert teleoperators on their first attempt (85.6% task progress / 80% success vs. 90.9% / 80.6% for experts with 375 h avg experience). Strong evidence that embodiment transfer is real and emergent.
The counterpoint: the AnyBody benchmark (arXiv 2505.14986) shows that truly novel-morphology extrapolation (new link structures, new compositions) still fails for every method tested. In-distribution cross-embodiment works; out-of-distribution morphology does not β yet.
This review maps the design space between those two poles.
flowchart TB
A[Cat A: Unified tokenized action space<br/>RT-X, Octo, RDT-1B, FAST]
B[Cat B: Soft-prompt / body-conditioning<br/>X-VLA, HPT]
C[Cat C: Embodiment-invariant latents<br/>UniVLA, UniAct, X-Sim, XR-1,<br/>UniSkill, TraceVLA]
D[Cat D: Human video / wearable bridging<br/>DexUMI, EgoDex, Visual-Imit-Humanoid,<br/>ImMimic, UniSkill]
E[Cat E: Morphology-aware architecture<br/>CrossFormer, HPT stems,<br/>GR00T N1, Octo, WholeBodyVLA]
F[Cat F: Scale + prompt expansion<br/>Ο0.5 β Ο0.6 β Ο0.7,<br/>Gemini Robotics 1.5]
G[Cat G: World-model-mediated<br/>DreamGen, Cosmos Policy, Ctrl-World]
CHOICE{Your situation}
CHOICE -- OXE-style data,<br/>simple baseline --> A
CHOICE -- many robots,<br/>cheap new-robot extension --> B
CHOICE -- humanβrobot transfer --> C
CHOICE -- data-starved --> D
CHOICE -- legs / wheels / wings --> E
CHOICE -- frontier lab data --> F
CHOICE -- no real data for target --> G
Β§5 details each category with exemplars, pros, cons, and when to pick it.
| Paper | Venue/year | Cat | Headline |
|---|---|---|---|
| Open X-Embodiment / RT-X | 2310.08864 / RSS 2024 | A | 22 robots unified via EE-pose tokens |
| Octo | 2405.12213 / CoRL 2024 | A+E | Diffusion head, modular I/O, fine-tune to new robot in hours |
| OpenVLA | CoRL 2024 | A | Open 7B on OXE |
| CrossFormer | 2408.11812 / CoRL 2024 | E | One transformer, 20 embodiments incl. quadcopter/quadruped |
| HPT | 2409.20537 / NeurIPS 2024 | B+E | Stem / trunk / head; 52 datasets |
| RDT-1B | 2410.07864 / ICLR 2025 | A | 1.2B diffusion, Physically Interpretable Unified Action Space |
| TraceVLA | 2412.10345 / ICLR 2025 | C | Visual trace overlay as embodiment-agnostic prompt |
| UniAct | 2501.10105 / CVPR 2025 | C+E | Universal atomic action codes; 28 embodiments |
| GR00T N1 / N1.5 / N1.6 | 2503.14734 / 2025 | E+F | Open humanoid foundation, dual-system VLA |
| Ο0.5 | 2504.16054 / CoRL 2025 Oral | F | Hierarchical + co-training; open-world homes |
| UniVLA (latent actions) | 2505.06111 / ICLR 2026 | C | Task-centric latent action tokens in DINO space |
| X-Sim | 2505.07096 / CoRL 2025 Oral | C+D | Real-to-sim-to-real via object motion |
| UniSkill | 2505.08787 / CoRL 2025 | C+D | Embodiment-invariant skill embedding from unlabeled video |
| DexUMI | 2505.21864 / CoRL 2025 Finalist | D | Wearable glove = universal manipulation interface |
| AnyBody (benchmark) | 2505.14986 | β | Shows novel-morphology extrapolation still fails |
| EgoDex | 2505.11709 / ICLR 2026 | D | 829h egocentric Vision Pro dex data |
| DreamGen | 2505.12705 / CoRL 2025 | G | Policy training inside video-world-model rollouts |
| ImMimic | 2509.10952 / CoRL 2025 Oral | D | Human-robot DTW + MixUp alignment; 4 gripper types |
| Visual Imit β Humanoid | CoRL 2025 Best Student | D+G | Internet video β sim humanoid β real humanoid |
| X-VLA | 2510.10274 / ICLR 2026 | B | Per-embodiment soft prompts, shared backbone, scaling laws |
| Gemini Robotics 1.5 | 2510.03342 / 2025 | F+C | Motion Transfer; ALOHA β bi-arm Franka β Apollo humanoid |
| XR-1 (UVMC) | ICLR 2026 | C | Shared vision-motion codebook |
| Cosmos Policy | ICLR 2026 | G+B | Cosmos video foundation backbone + control tokens |
| WholeBodyVLA | ICLR 2026 | E | Humanoid whole-body unified latent β coordinated base/arms |
| Ο0.6 | Nov 2025 | F | Gemma3-4B + KI + optional metadata |
| Ο0.7 | pi.website/pi07 / Apr 2026 | F | Metadata + MEM + subgoal images β zero-shot UR5e laundry β expert |
| Grasp2Grasp | NeurIPS 2025, 2506.02489 | C (grasp-level) | Princeton β SchrΓΆdinger-Bridge for cross-morphology grasp translation; stochastic transport between grasp distributions with physics costs |
- Mechanism: Unify 22 robots' data into one schema (~527 skills, ~160k tasks, 21 labs). RT-1-X and RT-2-X tokenize EE-pose deltas into discrete text tokens and train an LLM-style autoregressive policy on the union.
- Gap handled: Single-arm manipulators with different kinematics/cameras.
- Pros: Foundational dataset; positive transfer across 22 robots was a major 2023 result; the substrate every later VLA builds on.
- Cons: EE-pose-delta space is lowest-common-denominator; bimanual / dex / legged not well supported.
- Pipeline locus: Tokenizer (discrete EE-pose deltas).
- Mechanism: Transformer policy + diffusion action head; trained on 800k OXE trajectories to be a flexible initialization that fine-tunes to new sensory inputs and action spaces in hours.
- Pros: First truly open, downloadable cross-embodiment generalist; cheap fine-tune recipe.
- Cons: Per-robot specialists still beat it on hard tasks.
- Pipeline locus: Shared backbone + per-embodiment head swap.
- Mechanism: Single transformer with morphology-agnostic input handling β no hand-aligned obs/action formats. 900k trajectories across 20 embodiments including arms + wheeled base + quadcopter + quadruped.
- Pros: Matches specialists while spanning very different morphologies; first convincing "one net, many bodies" at the morphology level (not just many arms).
- Cons: Per-embodiment attention routing details modest; scaling behavior per-embodiment less characterized than X-VLA's.
- Pipeline locus: Unified tokenizer + shared trunk + per-embodiment readout heads.
- Mechanism: Stem / trunk / head factorization. Per-embodiment stems tokenize proprio + vision β fixed-size shared token sequence β shared trunk β per-task heads.
- Data: 52 datasets (real robots, sim, deployed, human video).
- Pros: Clean separation-of-concerns; +20% on unseen tasks; the precursor to X-VLA's "shared backbone + per-embodiment conditioning" idea.
- Cons: New robot = train a new stem (more machinery than X-VLA's single soft prompt).
- Pipeline locus: Embodiment-specific tokenizer (stem) + shared trunk + task head.
- Mechanism: Physically Interpretable Unified Action Space β explicit unified schema preserving physical meaning; 1.2B-param diffusion transformer specialized for bimanual manipulation.
- Pros: First serious 1B-scale diffusion foundation policy; strong bimanual; principled action-space unification.
- Cons: Unified action space lossy for dex hands; bimanual scope.
- Pipeline locus: Unified action space (training convention) + diffusion head.
- Mechanism: CoTracker extracts dense point trajectories from observation history; overlays active traces on the current frame; the VLA sees the overlay. Traces are embodiment-agnostic spatial-temporal context β they don't depend on arm appearance.
- Pros: +10% SimplerEnv, 3.5Γ real-robot vs. OpenVLA; no new parameters.
- Cons: 2D traces degrade under heavy occlusion; not an end-to-end universal action space.
- Pipeline locus: Input visual prompt (pixel-space).
- Mechanism: Vector-quantized vocabulary of universal atomic actions; a shared VLM outputs codes via Gumbel-Softmax; per-embodiment heads translate codes into robot-specific commands. 1M demos from 28 embodiments.
- Pros: A 0.5B UniAct reportedly outperforms 14Γ larger SOTA baselines on cross-embodiment evals. Explicit, interpretable atomic codes.
- Cons: Code count bounds expressiveness; fine-grained dex manipulation may not fit atomic codes cleanly.
- Pipeline locus: Intermediate representation (atomic action codes).
- Mechanism: Dual-system VLA β Eagle-2 VLM (System 2) + Diffusion Transformer (System 1). DiT blocks use an embodiment-aware state/action encoder that embeds robot-specific state + action.
- Data: Heterogeneous mix β real robot trajectories + human video + synthetic DreamGen rollouts.
- Pros: First open humanoid foundation model; tight sim/synthetic integration with NVIDIA's stack (Cosmos, Newton, DreamGen).
- Cons: Humanoid-leaning; embodiment-aware encoders still need per-robot calibration data.
- Pipeline locus: Embodiment-aware encoder inside DiT + VLM System 2.
- Mechanism: Hierarchical VLA: high-level subtask predictor + low-level flow-matching action expert. Co-training on multi-robot teleop + web VQA + object detection + semantic subtask prediction.
- Gap handled: Any arm the teleop mix covers (single-arm, bimanual, mobile).
- Pros: First open-world-home generalization at scale; template every subsequent generalist follows.
- Cons: Post-training often still needed for polish; no explicit cross-body mechanism β relies on scale + hierarchy.
- Pipeline locus: Scale + hierarchy.
- See: Ο0.5, Ο series evolution.
- Mechanism: Learn task-centric latent action tokens in DINO feature space via a latent-action model; humans + robots share this latent vocabulary; per-embodiment decoder maps latents β control.
- Pros: Beats OpenVLA with ~1/20 of pretrain compute and 1/10 of downstream data; works for navigation + manipulation.
- Cons: Bounded by DINO-feature quality; fine-grained contact may not fit the latent space; "task-centric" latents can collapse on visually busy, task-irrelevant video.
- Pipeline locus: Intermediate representation.
- Note: there are two ICLR 2026 papers called "UniVLA" β this is arXiv 2505.06111; the other is the Unified-VLA 8.5B at ICLR-2026-UniVLA.
- Mechanism: Extract object motion from human video β reproduce it in simulator β train image-conditioned diffusion policy β online sim-to-real adaptation. Sidesteps human-to-robot arm retargeting entirely.
- Pros: Principled; avoids the retargeting trap; clean recipe.
- Cons: Fails on tasks that aren't object-motion-reducible (force-sensitive insertion, gripper-geometry-critical); simulator fidelity is the ceiling.
- Pipeline locus: Intermediate representation (object motion) + sim-generated policy.
- See: X-Sim.
- Mechanism: Learn embodiment-agnostic skill embeddings from unlabeled cross-embodiment video (no aligned human-robot pairs needed); transfer to robot policies with few demos.
- Pros: Eliminates the expensive "aligned human+robot for same scene" data requirement.
- Cons: Encoder must separate skills visually; ambiguous actions with the same motion but different intent confuse the embedding.
- Pipeline locus: Intermediate representation (skill embedding).
- Mechanism: Wearable sensorized glove makes the human hand the universal manipulation interface. Record finger kinematics + contact β retarget to any robot hand.
- Pros: Breaks the dex-data bottleneck; captures natural grip diversity.
- Cons: Retargeting losses; some contact cues lost (palm, soft compliance); requires glove.
- Pipeline locus: Data-collection interface (front of pipeline).
- See: DexUMI.
- Mechanism: Not a method β a benchmark with 18 robot variations (8 procedurally generated + 10 real-world-based), two tasks (reach/push), with/without obstacles. Tests interpolation / extrapolation / composition over morphology.
- Findings: Zero-shot novel-morphology extrapolation and composition remain hard. In-distribution cross-embodiment works; genuinely new morphology does not.
- Why it matters: Canonical falsifier for "universal robot policy" claims. Every Group F method needs to confront this benchmark.
- Mechanism: 829 h of egocentric Vision Pro dex video, 194 tasks β a massive human-video dataset specifically designed to pretrain dexterous VLAs.
- Pros: Closes the data gap for Group D methods.
- Cons: Video only (no robot-action labels); retargeting is the user's problem.
- See: EgoDex.
- Mechanism: Train policies inside video-world-model rollouts; the world model abstracts embodiment dynamics at the pixel level.
- Pros: Infinite synthetic data; shared visual prior across robots.
- Cons: Pixel world models aren't yet precise enough for contact-heavy skills; heavy compute.
- Pipeline locus: World-model-mediated training data generator.
- See: DreamGen.
- Mechanism: Co-train on human video + small teleop set. DTW alignment maps retargeted human-hand poses to robot joints; MixUp interpolation creates intermediate humanβrobot domains.
- Pros: Explicitly builds a bridge rather than hoping transfer emerges. Works across 4 distinct end-effectors (Robotiq, Fin Ray, Allegro, Ability).
- Cons: Quality of retargeting + DTW alignment is the ceiling; MixUp in action space is a strong assumption.
- Pipeline locus: Data-level.
- Mechanism: Internet video β simulated humanoid β real humanoid. Pose estimation + retargeting + tracking in sim + distillation into a contextual whole-body policy.
- Pros: First credible "humanoid skills from internet video" at scale; activity-appropriate postures emerge.
- Cons: Limited by retargeter, sim fidelity, and human β humanoid mass-distribution differences.
- Pipeline locus: Data pipeline + sim training.
- See: Visual Imitation β Humanoid.
- Mechanism: Per-embodiment soft prompts (a handful of learned tokens per dataset) modulate a fully shared transformer backbone. New embodiment = train a new tiny prompt.
- Pros: Near-zero parameter cost per new robot; monotonic scaling law over # embodiments β the cleanest "scaling laws for cross-embodiment" story in 2026.
- Cons: Still needs demos per embodiment to fit the prompt; prompts are not interpretable; no explicit morphology topology.
- Pipeline locus: Prompt.
- See: X-VLA.
- Mechanism: Motion Transfer objective + novel architecture for learning from heterogeneous multi-embodiment data β zero-shot skill transfer across ALOHA β bi-arm Franka β Apollo humanoid without robot-specific post-training. Pairs with Gemini Robotics-ER 1.5/1.6 for embodied reasoning.
- Pros: Production-scale evidence of zero-shot cross-robot skill transfer; ALOHA β Apollo is a big morphology jump.
- Cons: Closed weights; mechanism details not fully public.
- Pipeline locus: Pretraining objective (Motion Transfer) + prompt/thinking stack.
- Mechanism: Dual-branch VQ-VAE (vision-dynamics branch + motion branch) with a shared codebook; downstream embodiments decode from the shared latent into their own action space.
- Pros: Strong inductive bias β forces an embodiment-free intermediate representation.
- Cons: Codebook quality is the bottleneck; fails if motion signals are too idiosyncratic (e.g., very different gripper kinematics); doesn't handle force/contact-only tasks.
- Pipeline locus: Tokenizer + shared codebook.
- See: XR-1.
- Mechanism: Fine-tune NVIDIA's Cosmos video foundation model with injected latent control tokens; model predicts next frame + action + value.
- Pros: Leverages industrial-scale video priors; interpretable future-frame predictions; supports value-based fine-tuning.
- Cons: Latents learned for generation, not control β precise low-level skills weaker; big inference cost.
- Pipeline locus: Backbone replacement (video foundation model) + token-level conditioning.
- See: Cosmos Policy.
- Mechanism: One VLA produces a unified latent that decodes to coordinated base / legs / arms / hands. Designed to capture loco β manip correlation.
- Gap handled: Humanoid whole-body (not fixed-base).
- Pros: First credible "generalist-VLA recipe" for whole-body humanoid.
- Cons: Humanoid-specific; not designed for non-humanoid transfer; paired data is scarce.
- Pipeline locus: Unified low-level decoder off a shared latent.
- See: WholeBodyVLA.
- Mechanism: Prompt expansion rather than architectural cross-embodiment magic. Subtask text + subgoal images (from a 14B BAGEL world model) + episode metadata (speed, quality 1β5, mistake flag) + control-mode token (joint vs. EE) + per-component dropout at train + CFG at inference. MEM video-history encoder handles temporal context.
- Gap handled: Anything the data mix covers β single-arm, bimanual, mobile, partial humanoid, UR5e.
- Headline: Zero-shot UR5e bimanual shirt-folding = 85.6% progress / 80% success, matching expert teleoperators (90.9% / 80.6%) on their first attempt with that robot.
- Data: Multi-robot teleop + web + DROID + egocentric human video + RL rollouts + failures + autonomous rollouts (disambiguated by metadata).
- Pros: State-of-the-art production system; one model β RL specialists; compositional generalization signal (unseen air fryer from 2 tangential episodes + web priors); discovers new grasp strategies on UR5e not present in training.
- Cons: Requires PI-scale data + metadata labeling; weights are proprietary; no head-to-head with ICLR 2026 discrete-diffusion VLAs.
- Pipeline locus: Prompt + data mix, zero architectural cross-embodiment magic. The UR5e result is arguably the strongest 2026 evidence that scale + prompting solves cross-embodiment for in-class morphologies.
- See: Ο0.7, Review-pi07, Ο series evolution.
ICRA 2026 (Vienna) adds a deployment-flavored layer to this map rather than a new paradigm; see the ICRA 2026 Survey. None of these claim novel-morphology extrapolation β consistent with AnyBody β and most stay firmly in-distribution.
-
OmniVLA (nav) (UC Berkeley, 2509.19480) extends Category F into navigation: an ~8.26B OpenVLA-class backbone trained on ~9,500 h across 10 platforms, with randomized modality fusion over pose / goal-image / language goals. Its goal-modality dropout is the navigation analogue of Ο0.7's prompt-component dropout, and it reports zero-shot cross-embodiment transfer to Unitree Go1 and VizBot β a rare in-the-wild data point for Cat F outside manipulation. (Name collision: a separate manipulation "OmniVLA," 2511.01210, fuses IR/radar/audio and is unrelated.)
-
Galaxea / G0 (2509.00576) is the cohort's counter-thesis to scale-driven Cat F: a 500 h / 100K-trajectory dataset on a single 23-DoF mobile-bimanual embodiment. Its three-stage curriculum cross-embodiment-pretrains (~1,000 h OXE) but finds the single-embodiment stage decisive, and in 20-trajectory few-shot tests embodiment-consistent pretraining beats no-pretraining. The takeaway sharpens Open Question Q1: raw cross-embodiment scale is not automatically better than consistent, well-annotated data for a fixed target.
-
Toward Embodiment-Equivariant VLA Policy (2509.14630) attacks cross-embodiment generalization through equivariance rather than scale β a structural inductive bias that sits alongside Category E's morphology-aware architectures and is, in spirit, the architecture-not-data reply AnyBody invites.
-
On the data/architecture side, OmniRetarget (2509.26633) is an interaction-preserving engine that augments a single demo across embodiments and terrains for humanoid loco-manipulation on a Unitree G1 β a Category D (retargeting) tool aimed at the humanoid whole-body regime β while DEFT (2602.02895) conditions a diffusion policy on embodiment and task so a damaged robot still completes it, a body-conditioning idea (Category B) repurposed for resilience rather than new-robot extension.
Net: ICRA 2026 mostly operationalizes existing categories (F omni-modal prompting, E equivariance, D retargeting, B body-conditioning) on real hardware, and Galaxea supplies the year's sharpest empirical caution against assuming more embodiments always helps.
- Mechanism: Discretize / bucket actions into one vocabulary β train LLM-style on everything.
- Exemplars: RT-X / OXE, OpenVLA, Octo, RDT-1B, FAST tokens / Ο0-FAST.
- Pros: β Simple Β· β leverages LLM training recipes Β· β scales with data.
- Cons: β Lossy for non-EE-pose actions (dex, humanoid) Β· β lowest-common-denominator.
- Pick when: You have OXE-style data and want a strong baseline fast.
- Mechanism: Tiny per-embodiment tokens modulate a fully shared backbone.
- Exemplars: X-VLA (canonical), HPT (stems are a heavier cousin).
- Pros: β Near-zero cost per new robot Β· β clean scaling laws over #embodiments Β· β backbone unchanged.
- Cons: β Still needs demos per embodiment to fit the prompt Β· β prompts not interpretable Β· β no morphology topology modeling.
- Pick when: You have many robots and want the cheapest possible extension to a new one.
- Mechanism: Learn a latent that isn't about motors: object motion, skill embedding, trace, atomic action code, or a grasp manifold.
- Exemplars: X-Sim (object motion), UniSkill (skill embedding), UniVLA-2505 (latent actions), UniAct (atomic codes), TraceVLA (visual traces), XR-1 (UVMC shared codebook). NeurIPS 2025 addition: Grasp2Grasp (Princeton, 2506.02489) β treats cross-morphology grasp transfer as stochastic transport between grasp distributions via SchrΓΆdinger Bridges, with physics-informed costs on base pose / contact map / wrench / manipulability. Applies specifically to grasp-level transfer (not full trajectory) but generalizes across very different hands.
- Pros: β Transfers across very different bodies Β· β strong inductive bias Β· β humanβrobot transfer works.
- Cons: β Latent quality is the ceiling Β· β fine-grained contact often doesn't fit Β· β ambiguous intents confuse latents.
- Pick when: You want clean humanβrobot or multi-morphology transfer.
- Mechanism: Humans are the universal demonstrators; bridge via retargeting, wearables, or alignment.
- Exemplars: DexUMI (glove), EgoDex (Vision Pro), Visual Imitation β Humanoid (internet video), ImMimic (DTW + MixUp), UniSkill (video skill embedding).
- Pros: β Unlocks massive unlabeled data Β· β natural behavior diversity Β· β no teleop bottleneck.
- Cons: β Retargeting loss Β· β force / contact often lost Β· β sim fidelity caps the ceiling.
- Pick when: You are data-starved and the task is visible in human video.
- Mechanism: Explicit per-robot input/output heads or routing.
- Exemplars: CrossFormer (20 embodiments incl. wheeled / aerial / legged), HPT (stem/trunk/head), GR00T N1 (embodiment-aware DiT encoder), Octo (diffusion head with modular I/O), WholeBodyVLA.
- Pros: β Respects per-robot physics Β· β handles very different morphologies (legs, wheels, wings) Β· β per-embodiment sensor configs are first-class.
- Cons: β More parameters per new robot Β· β harder to scale than Cat B.
- Pick when: Your "robots" include non-arm morphologies.
- Mechanism: Drown the gap in data; condition on rich prompts (subtask, subgoal image, metadata, control mode); let the model figure it out.
- Exemplars: Ο0.5 β Ο0.6 β Ο0.7, Gemini Robotics 1.5.
- Pros: β State-of-the-art in practice β Ο0.7's UR5e zero-shot = expert teleop Β· β minimal architectural baggage Β· β absorbs Categories A, C, D as prompt modalities.
- Cons: β Requires PI/Google-scale data & labeling discipline Β· β opaque about what's doing the work Β· β proprietary.
- Pick when: You have frontier-lab data access.
- Mechanism: Train policies inside a video world model that already abstracts embodiment.
- Exemplars: DreamGen, Cosmos Policy, Ctrl-World.
- Pros: β Infinite synthetic data Β· β shared visual prior across robots Β· β new-robot deployment without real data.
- Cons: β Pixel world models not yet precise enough for contact-heavy skills Β· β heavy compute Β· β sim-to-real still required.
- Pick when: You have no real data for the target embodiment.
| Axis | Best in class |
|---|---|
| Cheapest new-robot extension | X-VLA (B) β train a tiny soft prompt |
| Most diverse morphology span | CrossFormer (E) β arms + wheels + legs + wings |
| Most data-efficient transfer | UniVLA-2505 (C) β 1/20 compute, 1/10 data vs OpenVLA |
| Strongest zero-shot result in 2026 | Ο0.7 (F) β UR5e laundry β expert teleop |
| Most principled action-space unification | RDT-1B (A) Physically Interpretable UAS, UniAct (C) |
| Best for humanoid whole-body | WholeBodyVLA (E), GR00T N1.6 (E+F) |
| Best for dex hands with no teleop | DexUMI (D) |
| Best for "no real data, only video" | Visual Imitation β Humanoid (D+G), DreamGen (G) |
| Canonical falsifier of claims | AnyBody benchmark β shows novel-morphology still fails |
- 2023 β "can we even share?" RT-2 showed web VLMs help robots. Cross-embodiment was aspirational; teams hand-crafted per-robot stacks.
- Late 2023 / 2024 β Group A (unified tokens) dominates. OXE + RT-1-X / RT-2-X demonstrate positive transfer across 22 robots. OpenVLA opens the recipe. Working assumption: "tokenize EE-pose deltas, pretrain on everything."
- 2024 β Group E (morphology-aware architecture) appears. Octo, CrossFormer, HPT factor obs/action spaces by embodiment. The reply to RT-X: "we need architecture, not just data."
- Early-mid 2025 β Group C (invariant latents) blooms. UniAct atomic codes, UniVLA latent actions, TraceVLA visual traces, X-Sim object motion, UniSkill skill embeddings. Thesis: "don't unify the action β unify something that doesn't depend on the body."
- Mid 2025 β Group D (human video / wearable) mainstreams. DexUMI, EgoDex, Visual Imitation β Humanoid, ImMimic. Ο0.5 co-trains on web + human data. "The best cross-embodiment dataset is the human body."
- Late 2025 β Group F (scale + prompt expansion) crowns production. Ο0.5 β Ο0.6 (Nov 2025) β Ο0.7 (Apr 2026). Gemini Robotics 1.5. Thesis: "give the model enough data + rich prompts and it figures out the embodiment by itself." Ο0.7 zero-shot UR5e is the headline.
- Late 2025 / 2026 β ICLR 2026 synthesizes B + C + G. X-VLA = clean scaling laws via soft prompts; XR-1 = shared codebook across vision+motion; WholeBodyVLA = humanoid integration; DreamGen / Cosmos Policy / Ctrl-World = world models inside training.
Direction of travel:
- Mechanism consolidation. Soft-prompt (B) + invariant latent (C) + rich prompt (F) are converging β X-VLA's prompts and Ο0.7's metadata are the same idea from different angles.
- Data substrate unification. OXE + human video (EgoDex) + world-model rollouts (DreamGen / Cosmos) + RL traces (RECAP) is becoming the default pretraining mix.
- Benchmarks catch up. AnyBody (2025) makes novel-morphology generalization falsifiable β and shows we're not there yet on extrapolation/composition.
- Humanoid is in the tent. GR00T, WholeBodyVLA, Visual Imitation β Humanoid bring whole-body morphology into the same VLA discussion as tabletop.
Q1. Does scale alone solve cross-embodiment? Ο0.7's UR5e result suggests scale + prompting is enough for robots that "look like" your training distribution. But AnyBody shows extrapolation to new link structures still fails. ICRA 2026 adds a third caveat from the data side: Galaxea / G0 cross-embodiment-pretrains on ~1,000 h OXE but finds the single-embodiment stage decisive, and few-shot tests favor embodiment-consistent pretraining β i.e., raw cross-embodiment scale is not automatically better than consistent, well-annotated data for a fixed target. Updated answer: scale for coverage, architecture for extrapolation, and data consistency/quality for the deployment embodiment β three levers, not one. The "give it enough data and it figures out the body itself" thesis (trend #6) holds for in-class morphologies but is not a substitute for target-embodiment data quality.
Q2. Sim-to-real vs. real-to-sim for cross-embodiment?
- Sim-to-real (DreamGen, Cosmos Policy, Visual-Imit-Humanoid): best when task-completion is specifiable in sim and the embodiment gap is large (humanoid).
- Real-to-sim-to-real (X-Sim): best when human demos exist and tasks are object-motion-reducible.
- Real-only (Ο0.7, Gemini 1.5): best when frontier-lab data budgets are available.
- 2026 consensus: all three coexist; world-model-mediated sim (Group G) is closing the gap faster than expected.
Q3. Is there a universal action space that actually works?
- EE-pose-delta tokenized (RT-X, Octo) β works for single-arm, breaks on dex and humanoid.
- Physically Interpretable Unified Action Space (RDT-1B) β strictly better, still arm-centric.
- Atomic action codes (UniAct) β promising at coarse-skill level, unclear at contact-rich level.
- Latent action tokens (UniVLA-2505, XR-1 UVMC) β cleanest abstraction; tied to learned-feature quality.
- Empirical answer in 2026: no single representation wins across all skills. Production systems (Ο0.7) just condition on control-mode (joint vs EE) and let the flow-matching expert absorb the rest.
Q4. Can you transfer strategy, not just motion? Ο0.7 folding UR5e shirts reportedly discovers new grasp strategies not seen in training β strong strategy-transfer signal. UniSkill and UniAct also argue for strategy transfer via skill/atomic codes. But:
- Most cross-embodiment successes are within a class of strategies already in the data.
- Truly novel strategy invention (e.g., a parallel-jaw gripper discovering a pinch it has never seen) remains rare.
- The Ο0.7 UR5e result is the strongest 2026 evidence of genuine strategy transfer; academic reproducibility is the key question for 2026β2027.
Q5. Are benchmarks keeping up? AnyBody (2025), RoboArena β, RoboCasa365, WorldGym all push this direction. Expect a shakeout in 2026 where Group F claims get tested on AnyBody-style extrapolation suites β and we see which ones survive.
Q6. Closed vs. open ecosystems. The strongest cross-embodiment systems (Ο0.7, Gemini Robotics 1.5) are closed-weights. The strongest open alternatives (GR00T N1.6, OpenVLA, Octo, RDT-1B) trail on the hardest tasks. Whether open ecosystems close this gap in 2026 depends on who ships the next large multi-embodiment dataset β and whether it comes from academia (OXE v2?), industry (NVIDIA?), or a coalition.
If you're shipping a cross-embodiment VLA and wondering where to spend a quarter:
- OXE-style data on a few single-arm robots? β Start with Octo / OpenVLA / RDT-1B as baselines (Cat A). Don't overthink it.
- Many robots, want cheapest per-new-robot cost? β X-VLA-style soft prompts on a shared backbone (Cat B).
- Big human-robot gap or human-video-abundant task? β Pick an invariant-latent method (Cat C β UniAct or X-Sim), or bridge via human demos (Cat D β DexUMI, EgoDex, ImMimic).
- Legs / wheels / wings in the mix? β Go morphology-aware (Cat E β CrossFormer or GR00T N1).
- Humanoid whole-body target? β WholeBodyVLA (Cat E) or GR00T N1.6 (Cat E+F) with paired loco-manipulation data.
- Have frontier-scale data budget? β Follow the Ο0.7 / Gemini 1.5 playbook (Cat F): prompt expansion + massive multi-robot mix + RL rollouts distilled back in.
- No real data for the target embodiment? β DreamGen or Cosmos Policy rollouts (Cat G) as a bootstrap; fine-tune on the real target once you have any data.
Don't skip benchmarking on AnyBody-style extrapolation. If your system only works on robots that look like the training distribution, that's fine β but the deployment story differs a lot from "universal."
- Ο series: Ο0.5 Β· Ο0.6 Β· Ο0.7 Β· Ο series evolution Β· Review-pi07
- ICLR 2026 cross-embodiment cluster: X-VLA Β· XR-1 Β· UniVLA Β· Cosmos Policy Β· WholeBodyVLA Β· Ctrl-World
- CoRL 2025 cross-embodiment cluster: X-Sim Β· DexUMI Β· Visual Imitation β Humanoid Β· DreamGen
- Data pages: EgoDex Β· RoboCasa365
- Companion reviews: Review-VLA-Memory Β· Review-pi07 Β· Review-VLM4VLA
- Surveys: ICLR 2026 survey Β· CoRL 2025 survey
Verdict: 2026 flipped the question β action-representation alignment is the precondition for cross-embodiment data to scale at all; hands joined the agenda with 80%+ zero-shot transfer.
Qwen-RobotManip's controlled scaling ablation is the pivot: naive zero-padded spaces yield no scaling law; a canonical 80-dim representation scales; camera-frame delta EEF scales log-linearly and transfers zero-shot to unseen arms (RoboTwin-XE 23.9% vs 14.5% joint; UR5 5.6Γ better). RSS 2026 extended to hands: DexGrasp-Zero (morphology-aligned graphs + fixed per-hand mapping, 85% zero-shot, +59.5% over SOTA) and One-Hand (canonical URDF + smooth morphology latent, 81.9%); LAP-style language-action pretraining and X-DiffVLA-class cross-embodied heads round out the cluster.
| Interface | Pros | Cons | Proven at |
|---|---|---|---|
| Canonical padded tensor + masks | Simple, scales, one policy | Slot semantics curated per embodiment | Qwen suite, GR00T-class |
| Camera-frame delta EEF (+CaPE) | Visual-action alignment; best zero-shot; creates the scaling law | Calibrated cameras needed at train & test | RobotManip |
| Morphology graphs / canonical URDFs | Anatomy-grounded; no trainable retargeting | Hands-only so far | DexGrasp-Zero, One-Hand |
| Embodiment prompts + in-context history | No architecture change | Weak alone (+2β3 pp; history costs denoise steps) | RobotManip ablations |
| Latent action/motion codes (ICML 2026) | Embodiment-agnostic by construction | Decoding fidelity per platform | XR-1 (Oral: unified vision-motion codes, 12k+ rollouts / 6 embodiments); LAC-WM (+46.7% unseen-robot WM adaptation) |
| Data-side augmentation (ICML 2026) | Works with any interface | Synthetic-render gap | OXE-AugE re-renders OXE across 9 embodiments (4.4M trajectories; +24β45% on unseen robot-gripper combos); RDT2 scales 10k+ h UMI data to zero-shot cross-embodiment |
- Joint-space zero-shot transfer still fails (<5% on unseen morphologies).
- Morphology distance matters: Franka-class 7-DoF gaps resist even camera-frame transfer (5.9%); gripperβdexterous-hand untried.
- Calibration is a hidden deployment tax for the best interface; uncalibrated fallbacks under-evaluated.
- Humanβrobot is the hardest pair β see Review-Human-Video-Transfer.
β Back to Home