Review Cross Embodiment - Heungwoo/research GitHub Wiki

In-Depth Review β€” Cross-Embodiment Training for VLAs

Compiled April 2026 Β· Focus: how recent VLA papers (Ο€-series + ICLR 2026 + CoRL 2025 + 2024–2026 arXiv) train one policy for many bodies, and how those designs compare.

This is a cross-paper review page organized by how the training data is mixed. For the deployment cut β€” which papers run one frozen checkpoint on multiple robots at inference (routed seen-robot vs zero-shot-to-unseen) β€” see the companion Single-Checkpoint Multi-Robot Deployment. For per-paper summaries see the linked Cross-embodiment pages. Related reviews: Review-VLA-Memory, Review-pi07, Review-VLM4VLA.


1. Why cross-embodiment matters

Every robot lab collects data on its own hardware. Every VLA needs to run on the hardware the customer has. The gap between those two is the cross-embodiment problem:

  • Arm / gripper differences β€” different reach, DoF, force envelope, end-effector geometry.
  • Base differences β€” fixed vs. wheeled vs. legged vs. humanoid, changes the whole workspace.
  • Sensor differences β€” wrist cam vs. head cam; 3 views vs. 4.
  • Control differences β€” joint space vs. end-effector space; 20 Hz vs. 50 Hz.
  • Strategy differences β€” a bigger / heavier arm may need a completely different grasp strategy for the same task.

The headline 2026 result: Ο€0.7 zero-shots laundry folding on a UR5e bimanual that was never in its laundry training data, matching expert teleoperators on their first attempt (85.6% task progress / 80% success vs. 90.9% / 80.6% for experts with 375 h avg experience). Strong evidence that embodiment transfer is real and emergent.

The counterpoint: the AnyBody benchmark (arXiv 2505.14986) shows that truly novel-morphology extrapolation (new link structures, new compositions) still fails for every method tested. In-distribution cross-embodiment works; out-of-distribution morphology does not β€” yet.

This review maps the design space between those two poles.


2. The seven categories β€” at a glance

flowchart TB
  A[Cat A: Unified tokenized action space<br/>RT-X, Octo, RDT-1B, FAST]
  B[Cat B: Soft-prompt / body-conditioning<br/>X-VLA, HPT]
  C[Cat C: Embodiment-invariant latents<br/>UniVLA, UniAct, X-Sim, XR-1,<br/>UniSkill, TraceVLA]
  D[Cat D: Human video / wearable bridging<br/>DexUMI, EgoDex, Visual-Imit-Humanoid,<br/>ImMimic, UniSkill]
  E[Cat E: Morphology-aware architecture<br/>CrossFormer, HPT stems,<br/>GR00T N1, Octo, WholeBodyVLA]
  F[Cat F: Scale + prompt expansion<br/>Ο€0.5 β†’ Ο€0.6 β†’ Ο€0.7,<br/>Gemini Robotics 1.5]
  G[Cat G: World-model-mediated<br/>DreamGen, Cosmos Policy, Ctrl-World]

  CHOICE{Your situation}
  CHOICE -- OXE-style data,<br/>simple baseline --> A
  CHOICE -- many robots,<br/>cheap new-robot extension --> B
  CHOICE -- human↔robot transfer --> C
  CHOICE -- data-starved --> D
  CHOICE -- legs / wheels / wings --> E
  CHOICE -- frontier lab data --> F
  CHOICE -- no real data for target --> G
Loading

Β§5 details each category with exemplars, pros, cons, and when to pick it.


3. The papers in this review

Paper Venue/year Cat Headline
Open X-Embodiment / RT-X 2310.08864 / RSS 2024 A 22 robots unified via EE-pose tokens
Octo 2405.12213 / CoRL 2024 A+E Diffusion head, modular I/O, fine-tune to new robot in hours
OpenVLA CoRL 2024 A Open 7B on OXE
CrossFormer 2408.11812 / CoRL 2024 E One transformer, 20 embodiments incl. quadcopter/quadruped
HPT 2409.20537 / NeurIPS 2024 B+E Stem / trunk / head; 52 datasets
RDT-1B 2410.07864 / ICLR 2025 A 1.2B diffusion, Physically Interpretable Unified Action Space
TraceVLA 2412.10345 / ICLR 2025 C Visual trace overlay as embodiment-agnostic prompt
UniAct 2501.10105 / CVPR 2025 C+E Universal atomic action codes; 28 embodiments
GR00T N1 / N1.5 / N1.6 2503.14734 / 2025 E+F Open humanoid foundation, dual-system VLA
Ο€0.5 2504.16054 / CoRL 2025 Oral F Hierarchical + co-training; open-world homes
UniVLA (latent actions) 2505.06111 / ICLR 2026 C Task-centric latent action tokens in DINO space
X-Sim 2505.07096 / CoRL 2025 Oral C+D Real-to-sim-to-real via object motion
UniSkill 2505.08787 / CoRL 2025 C+D Embodiment-invariant skill embedding from unlabeled video
DexUMI 2505.21864 / CoRL 2025 Finalist D Wearable glove = universal manipulation interface
AnyBody (benchmark) 2505.14986 β€” Shows novel-morphology extrapolation still fails
EgoDex 2505.11709 / ICLR 2026 D 829h egocentric Vision Pro dex data
DreamGen 2505.12705 / CoRL 2025 G Policy training inside video-world-model rollouts
ImMimic 2509.10952 / CoRL 2025 Oral D Human-robot DTW + MixUp alignment; 4 gripper types
Visual Imit β†’ Humanoid CoRL 2025 Best Student D+G Internet video β†’ sim humanoid β†’ real humanoid
X-VLA 2510.10274 / ICLR 2026 B Per-embodiment soft prompts, shared backbone, scaling laws
Gemini Robotics 1.5 2510.03342 / 2025 F+C Motion Transfer; ALOHA ↔ bi-arm Franka ↔ Apollo humanoid
XR-1 (UVMC) ICLR 2026 C Shared vision-motion codebook
Cosmos Policy ICLR 2026 G+B Cosmos video foundation backbone + control tokens
WholeBodyVLA ICLR 2026 E Humanoid whole-body unified latent β†’ coordinated base/arms
Ο€0.6 Nov 2025 F Gemma3-4B + KI + optional metadata
Ο€0.7 pi.website/pi07 / Apr 2026 F Metadata + MEM + subgoal images β†’ zero-shot UR5e laundry β‰ˆ expert
Grasp2Grasp NeurIPS 2025, 2506.02489 C (grasp-level) Princeton β€” SchrΓΆdinger-Bridge for cross-morphology grasp translation; stochastic transport between grasp distributions with physics costs

4. Deep-dives

4.1 Open X-Embodiment & RT-X (2023, RSS/Science 2024)

  • Mechanism: Unify 22 robots' data into one schema (~527 skills, ~160k tasks, 21 labs). RT-1-X and RT-2-X tokenize EE-pose deltas into discrete text tokens and train an LLM-style autoregressive policy on the union.
  • Gap handled: Single-arm manipulators with different kinematics/cameras.
  • Pros: Foundational dataset; positive transfer across 22 robots was a major 2023 result; the substrate every later VLA builds on.
  • Cons: EE-pose-delta space is lowest-common-denominator; bimanual / dex / legged not well supported.
  • Pipeline locus: Tokenizer (discrete EE-pose deltas).

4.2 Octo (Ghosh et al., ICRA/CoRL 2024)

  • Mechanism: Transformer policy + diffusion action head; trained on 800k OXE trajectories to be a flexible initialization that fine-tunes to new sensory inputs and action spaces in hours.
  • Pros: First truly open, downloadable cross-embodiment generalist; cheap fine-tune recipe.
  • Cons: Per-robot specialists still beat it on hard tasks.
  • Pipeline locus: Shared backbone + per-embodiment head swap.

4.3 CrossFormer (CoRL 2024)

  • Mechanism: Single transformer with morphology-agnostic input handling β€” no hand-aligned obs/action formats. 900k trajectories across 20 embodiments including arms + wheeled base + quadcopter + quadruped.
  • Pros: Matches specialists while spanning very different morphologies; first convincing "one net, many bodies" at the morphology level (not just many arms).
  • Cons: Per-embodiment attention routing details modest; scaling behavior per-embodiment less characterized than X-VLA's.
  • Pipeline locus: Unified tokenizer + shared trunk + per-embodiment readout heads.

4.4 HPT β€” Heterogeneous Pretrained Transformers (NeurIPS 2024)

  • Mechanism: Stem / trunk / head factorization. Per-embodiment stems tokenize proprio + vision β†’ fixed-size shared token sequence β†’ shared trunk β†’ per-task heads.
  • Data: 52 datasets (real robots, sim, deployed, human video).
  • Pros: Clean separation-of-concerns; +20% on unseen tasks; the precursor to X-VLA's "shared backbone + per-embodiment conditioning" idea.
  • Cons: New robot = train a new stem (more machinery than X-VLA's single soft prompt).
  • Pipeline locus: Embodiment-specific tokenizer (stem) + shared trunk + task head.

4.5 RDT-1B (ICLR 2025)

  • Mechanism: Physically Interpretable Unified Action Space β€” explicit unified schema preserving physical meaning; 1.2B-param diffusion transformer specialized for bimanual manipulation.
  • Pros: First serious 1B-scale diffusion foundation policy; strong bimanual; principled action-space unification.
  • Cons: Unified action space lossy for dex hands; bimanual scope.
  • Pipeline locus: Unified action space (training convention) + diffusion head.

4.6 TraceVLA (ICLR 2025)

  • Mechanism: CoTracker extracts dense point trajectories from observation history; overlays active traces on the current frame; the VLA sees the overlay. Traces are embodiment-agnostic spatial-temporal context β€” they don't depend on arm appearance.
  • Pros: +10% SimplerEnv, 3.5Γ— real-robot vs. OpenVLA; no new parameters.
  • Cons: 2D traces degrade under heavy occlusion; not an end-to-end universal action space.
  • Pipeline locus: Input visual prompt (pixel-space).

4.7 UniAct β€” Universal Atomic Actions (CVPR 2025)

  • Mechanism: Vector-quantized vocabulary of universal atomic actions; a shared VLM outputs codes via Gumbel-Softmax; per-embodiment heads translate codes into robot-specific commands. 1M demos from 28 embodiments.
  • Pros: A 0.5B UniAct reportedly outperforms 14Γ— larger SOTA baselines on cross-embodiment evals. Explicit, interpretable atomic codes.
  • Cons: Code count bounds expressiveness; fine-grained dex manipulation may not fit atomic codes cleanly.
  • Pipeline locus: Intermediate representation (atomic action codes).

4.8 NVIDIA Isaac GR00T N1 / N1.5 / N1.6 (2025)

  • Mechanism: Dual-system VLA β€” Eagle-2 VLM (System 2) + Diffusion Transformer (System 1). DiT blocks use an embodiment-aware state/action encoder that embeds robot-specific state + action.
  • Data: Heterogeneous mix β€” real robot trajectories + human video + synthetic DreamGen rollouts.
  • Pros: First open humanoid foundation model; tight sim/synthetic integration with NVIDIA's stack (Cosmos, Newton, DreamGen).
  • Cons: Humanoid-leaning; embodiment-aware encoders still need per-robot calibration data.
  • Pipeline locus: Embodiment-aware encoder inside DiT + VLM System 2.

4.9 Ο€0.5 (CoRL 2025 Oral)

  • Mechanism: Hierarchical VLA: high-level subtask predictor + low-level flow-matching action expert. Co-training on multi-robot teleop + web VQA + object detection + semantic subtask prediction.
  • Gap handled: Any arm the teleop mix covers (single-arm, bimanual, mobile).
  • Pros: First open-world-home generalization at scale; template every subsequent generalist follows.
  • Cons: Post-training often still needed for polish; no explicit cross-body mechanism β€” relies on scale + hierarchy.
  • Pipeline locus: Scale + hierarchy.
  • See: Ο€0.5, Ο€ series evolution.

4.10 UniVLA β€” Latent Action Tokens (ICLR 2026, arXiv 2505.06111)

  • Mechanism: Learn task-centric latent action tokens in DINO feature space via a latent-action model; humans + robots share this latent vocabulary; per-embodiment decoder maps latents β†’ control.
  • Pros: Beats OpenVLA with ~1/20 of pretrain compute and 1/10 of downstream data; works for navigation + manipulation.
  • Cons: Bounded by DINO-feature quality; fine-grained contact may not fit the latent space; "task-centric" latents can collapse on visually busy, task-irrelevant video.
  • Pipeline locus: Intermediate representation.
  • Note: there are two ICLR 2026 papers called "UniVLA" β€” this is arXiv 2505.06111; the other is the Unified-VLA 8.5B at ICLR-2026-UniVLA.

4.11 X-Sim (CoRL 2025 Oral)

  • Mechanism: Extract object motion from human video β†’ reproduce it in simulator β†’ train image-conditioned diffusion policy β†’ online sim-to-real adaptation. Sidesteps human-to-robot arm retargeting entirely.
  • Pros: Principled; avoids the retargeting trap; clean recipe.
  • Cons: Fails on tasks that aren't object-motion-reducible (force-sensitive insertion, gripper-geometry-critical); simulator fidelity is the ceiling.
  • Pipeline locus: Intermediate representation (object motion) + sim-generated policy.
  • See: X-Sim.

4.12 UniSkill (CoRL 2025)

  • Mechanism: Learn embodiment-agnostic skill embeddings from unlabeled cross-embodiment video (no aligned human-robot pairs needed); transfer to robot policies with few demos.
  • Pros: Eliminates the expensive "aligned human+robot for same scene" data requirement.
  • Cons: Encoder must separate skills visually; ambiguous actions with the same motion but different intent confuse the embedding.
  • Pipeline locus: Intermediate representation (skill embedding).

4.13 DexUMI (CoRL 2025 Finalist)

  • Mechanism: Wearable sensorized glove makes the human hand the universal manipulation interface. Record finger kinematics + contact β†’ retarget to any robot hand.
  • Pros: Breaks the dex-data bottleneck; captures natural grip diversity.
  • Cons: Retargeting losses; some contact cues lost (palm, soft compliance); requires glove.
  • Pipeline locus: Data-collection interface (front of pipeline).
  • See: DexUMI.

4.14 AnyBody (benchmark, arXiv 2505.14986)

  • Mechanism: Not a method β€” a benchmark with 18 robot variations (8 procedurally generated + 10 real-world-based), two tasks (reach/push), with/without obstacles. Tests interpolation / extrapolation / composition over morphology.
  • Findings: Zero-shot novel-morphology extrapolation and composition remain hard. In-distribution cross-embodiment works; genuinely new morphology does not.
  • Why it matters: Canonical falsifier for "universal robot policy" claims. Every Group F method needs to confront this benchmark.

4.15 EgoDex (ICLR 2026)

  • Mechanism: 829 h of egocentric Vision Pro dex video, 194 tasks β€” a massive human-video dataset specifically designed to pretrain dexterous VLAs.
  • Pros: Closes the data gap for Group D methods.
  • Cons: Video only (no robot-action labels); retargeting is the user's problem.
  • See: EgoDex.

4.16 DreamGen (CoRL 2025)

  • Mechanism: Train policies inside video-world-model rollouts; the world model abstracts embodiment dynamics at the pixel level.
  • Pros: Infinite synthetic data; shared visual prior across robots.
  • Cons: Pixel world models aren't yet precise enough for contact-heavy skills; heavy compute.
  • Pipeline locus: World-model-mediated training data generator.
  • See: DreamGen.

4.17 ImMimic (CoRL 2025 Oral)

  • Mechanism: Co-train on human video + small teleop set. DTW alignment maps retargeted human-hand poses to robot joints; MixUp interpolation creates intermediate human↔robot domains.
  • Pros: Explicitly builds a bridge rather than hoping transfer emerges. Works across 4 distinct end-effectors (Robotiq, Fin Ray, Allegro, Ability).
  • Cons: Quality of retargeting + DTW alignment is the ceiling; MixUp in action space is a strong assumption.
  • Pipeline locus: Data-level.

4.18 Visual Imitation β†’ Humanoid (CoRL 2025 Best Student)

  • Mechanism: Internet video β†’ simulated humanoid β†’ real humanoid. Pose estimation + retargeting + tracking in sim + distillation into a contextual whole-body policy.
  • Pros: First credible "humanoid skills from internet video" at scale; activity-appropriate postures emerge.
  • Cons: Limited by retargeter, sim fidelity, and human ↔ humanoid mass-distribution differences.
  • Pipeline locus: Data pipeline + sim training.
  • See: Visual Imitation β†’ Humanoid.

4.19 X-VLA (ICLR 2026)

  • Mechanism: Per-embodiment soft prompts (a handful of learned tokens per dataset) modulate a fully shared transformer backbone. New embodiment = train a new tiny prompt.
  • Pros: Near-zero parameter cost per new robot; monotonic scaling law over # embodiments β€” the cleanest "scaling laws for cross-embodiment" story in 2026.
  • Cons: Still needs demos per embodiment to fit the prompt; prompts are not interpretable; no explicit morphology topology.
  • Pipeline locus: Prompt.
  • See: X-VLA.

4.20 Gemini Robotics 1.5 (DeepMind 2025)

  • Mechanism: Motion Transfer objective + novel architecture for learning from heterogeneous multi-embodiment data β†’ zero-shot skill transfer across ALOHA ↔ bi-arm Franka ↔ Apollo humanoid without robot-specific post-training. Pairs with Gemini Robotics-ER 1.5/1.6 for embodied reasoning.
  • Pros: Production-scale evidence of zero-shot cross-robot skill transfer; ALOHA β†’ Apollo is a big morphology jump.
  • Cons: Closed weights; mechanism details not fully public.
  • Pipeline locus: Pretraining objective (Motion Transfer) + prompt/thinking stack.

4.21 XR-1 / Unified Vision-Motion Codes (ICLR 2026)

  • Mechanism: Dual-branch VQ-VAE (vision-dynamics branch + motion branch) with a shared codebook; downstream embodiments decode from the shared latent into their own action space.
  • Pros: Strong inductive bias β€” forces an embodiment-free intermediate representation.
  • Cons: Codebook quality is the bottleneck; fails if motion signals are too idiosyncratic (e.g., very different gripper kinematics); doesn't handle force/contact-only tasks.
  • Pipeline locus: Tokenizer + shared codebook.
  • See: XR-1.

4.22 Cosmos Policy (ICLR 2026)

  • Mechanism: Fine-tune NVIDIA's Cosmos video foundation model with injected latent control tokens; model predicts next frame + action + value.
  • Pros: Leverages industrial-scale video priors; interpretable future-frame predictions; supports value-based fine-tuning.
  • Cons: Latents learned for generation, not control β€” precise low-level skills weaker; big inference cost.
  • Pipeline locus: Backbone replacement (video foundation model) + token-level conditioning.
  • See: Cosmos Policy.

4.23 WholeBodyVLA (ICLR 2026)

  • Mechanism: One VLA produces a unified latent that decodes to coordinated base / legs / arms / hands. Designed to capture loco ↔ manip correlation.
  • Gap handled: Humanoid whole-body (not fixed-base).
  • Pros: First credible "generalist-VLA recipe" for whole-body humanoid.
  • Cons: Humanoid-specific; not designed for non-humanoid transfer; paired data is scarce.
  • Pipeline locus: Unified low-level decoder off a shared latent.
  • See: WholeBodyVLA.

4.24 Ο€0.7 (Apr 2026) β€” the 2026 benchmark

  • Mechanism: Prompt expansion rather than architectural cross-embodiment magic. Subtask text + subgoal images (from a 14B BAGEL world model) + episode metadata (speed, quality 1–5, mistake flag) + control-mode token (joint vs. EE) + per-component dropout at train + CFG at inference. MEM video-history encoder handles temporal context.
  • Gap handled: Anything the data mix covers β€” single-arm, bimanual, mobile, partial humanoid, UR5e.
  • Headline: Zero-shot UR5e bimanual shirt-folding = 85.6% progress / 80% success, matching expert teleoperators (90.9% / 80.6%) on their first attempt with that robot.
  • Data: Multi-robot teleop + web + DROID + egocentric human video + RL rollouts + failures + autonomous rollouts (disambiguated by metadata).
  • Pros: State-of-the-art production system; one model β‰ˆ RL specialists; compositional generalization signal (unseen air fryer from 2 tangential episodes + web priors); discovers new grasp strategies on UR5e not present in training.
  • Cons: Requires PI-scale data + metadata labeling; weights are proprietary; no head-to-head with ICLR 2026 discrete-diffusion VLAs.
  • Pipeline locus: Prompt + data mix, zero architectural cross-embodiment magic. The UR5e result is arguably the strongest 2026 evidence that scale + prompting solves cross-embodiment for in-class morphologies.
  • See: Ο€0.7, Review-pi07, Ο€ series evolution.

4b. ICRA 2026 developments

ICRA 2026 (Vienna) adds a deployment-flavored layer to this map rather than a new paradigm; see the ICRA 2026 Survey. None of these claim novel-morphology extrapolation β€” consistent with AnyBody β€” and most stay firmly in-distribution.

  • OmniVLA (nav) (UC Berkeley, 2509.19480) extends Category F into navigation: an ~8.26B OpenVLA-class backbone trained on ~9,500 h across 10 platforms, with randomized modality fusion over pose / goal-image / language goals. Its goal-modality dropout is the navigation analogue of Ο€0.7's prompt-component dropout, and it reports zero-shot cross-embodiment transfer to Unitree Go1 and VizBot β€” a rare in-the-wild data point for Cat F outside manipulation. (Name collision: a separate manipulation "OmniVLA," 2511.01210, fuses IR/radar/audio and is unrelated.)

  • Galaxea / G0 (2509.00576) is the cohort's counter-thesis to scale-driven Cat F: a 500 h / 100K-trajectory dataset on a single 23-DoF mobile-bimanual embodiment. Its three-stage curriculum cross-embodiment-pretrains (~1,000 h OXE) but finds the single-embodiment stage decisive, and in 20-trajectory few-shot tests embodiment-consistent pretraining beats no-pretraining. The takeaway sharpens Open Question Q1: raw cross-embodiment scale is not automatically better than consistent, well-annotated data for a fixed target.

  • Toward Embodiment-Equivariant VLA Policy (2509.14630) attacks cross-embodiment generalization through equivariance rather than scale β€” a structural inductive bias that sits alongside Category E's morphology-aware architectures and is, in spirit, the architecture-not-data reply AnyBody invites.

  • On the data/architecture side, OmniRetarget (2509.26633) is an interaction-preserving engine that augments a single demo across embodiments and terrains for humanoid loco-manipulation on a Unitree G1 β€” a Category D (retargeting) tool aimed at the humanoid whole-body regime β€” while DEFT (2602.02895) conditions a diffusion policy on embodiment and task so a damaged robot still completes it, a body-conditioning idea (Category B) repurposed for resilience rather than new-robot extension.

Net: ICRA 2026 mostly operationalizes existing categories (F omni-modal prompting, E equivariance, D retargeting, B body-conditioning) on real hardware, and Galaxea supplies the year's sharpest empirical caution against assuming more embodiments always helps.


5. Categories compared β€” pros, cons, when to pick

Category A β€” Unified tokenized action space

  • Mechanism: Discretize / bucket actions into one vocabulary β†’ train LLM-style on everything.
  • Exemplars: RT-X / OXE, OpenVLA, Octo, RDT-1B, FAST tokens / Ο€0-FAST.
  • Pros: βœ… Simple Β· βœ… leverages LLM training recipes Β· βœ… scales with data.
  • Cons: ❌ Lossy for non-EE-pose actions (dex, humanoid) Β· ❌ lowest-common-denominator.
  • Pick when: You have OXE-style data and want a strong baseline fast.

Category B β€” Soft-prompt / body-conditioning

  • Mechanism: Tiny per-embodiment tokens modulate a fully shared backbone.
  • Exemplars: X-VLA (canonical), HPT (stems are a heavier cousin).
  • Pros: βœ… Near-zero cost per new robot Β· βœ… clean scaling laws over #embodiments Β· βœ… backbone unchanged.
  • Cons: ❌ Still needs demos per embodiment to fit the prompt Β· ❌ prompts not interpretable Β· ❌ no morphology topology modeling.
  • Pick when: You have many robots and want the cheapest possible extension to a new one.

Category C β€” Embodiment-invariant intermediate representation

  • Mechanism: Learn a latent that isn't about motors: object motion, skill embedding, trace, atomic action code, or a grasp manifold.
  • Exemplars: X-Sim (object motion), UniSkill (skill embedding), UniVLA-2505 (latent actions), UniAct (atomic codes), TraceVLA (visual traces), XR-1 (UVMC shared codebook). NeurIPS 2025 addition: Grasp2Grasp (Princeton, 2506.02489) β€” treats cross-morphology grasp transfer as stochastic transport between grasp distributions via SchrΓΆdinger Bridges, with physics-informed costs on base pose / contact map / wrench / manipulability. Applies specifically to grasp-level transfer (not full trajectory) but generalizes across very different hands.
  • Pros: βœ… Transfers across very different bodies Β· βœ… strong inductive bias Β· βœ… human↔robot transfer works.
  • Cons: ❌ Latent quality is the ceiling Β· ❌ fine-grained contact often doesn't fit Β· ❌ ambiguous intents confuse latents.
  • Pick when: You want clean human↔robot or multi-morphology transfer.

Category D β€” Human video / wearable bridging

  • Mechanism: Humans are the universal demonstrators; bridge via retargeting, wearables, or alignment.
  • Exemplars: DexUMI (glove), EgoDex (Vision Pro), Visual Imitation β†’ Humanoid (internet video), ImMimic (DTW + MixUp), UniSkill (video skill embedding).
  • Pros: βœ… Unlocks massive unlabeled data Β· βœ… natural behavior diversity Β· βœ… no teleop bottleneck.
  • Cons: ❌ Retargeting loss Β· ❌ force / contact often lost Β· ❌ sim fidelity caps the ceiling.
  • Pick when: You are data-starved and the task is visible in human video.

Category E β€” Morphology-aware architecture

  • Mechanism: Explicit per-robot input/output heads or routing.
  • Exemplars: CrossFormer (20 embodiments incl. wheeled / aerial / legged), HPT (stem/trunk/head), GR00T N1 (embodiment-aware DiT encoder), Octo (diffusion head with modular I/O), WholeBodyVLA.
  • Pros: βœ… Respects per-robot physics Β· βœ… handles very different morphologies (legs, wheels, wings) Β· βœ… per-embodiment sensor configs are first-class.
  • Cons: ❌ More parameters per new robot Β· ❌ harder to scale than Cat B.
  • Pick when: Your "robots" include non-arm morphologies.

Category F β€” Scale + prompt expansion

  • Mechanism: Drown the gap in data; condition on rich prompts (subtask, subgoal image, metadata, control mode); let the model figure it out.
  • Exemplars: Ο€0.5 β†’ Ο€0.6 β†’ Ο€0.7, Gemini Robotics 1.5.
  • Pros: βœ… State-of-the-art in practice β€” Ο€0.7's UR5e zero-shot = expert teleop Β· βœ… minimal architectural baggage Β· βœ… absorbs Categories A, C, D as prompt modalities.
  • Cons: ❌ Requires PI/Google-scale data & labeling discipline Β· ❌ opaque about what's doing the work Β· ❌ proprietary.
  • Pick when: You have frontier-lab data access.

Category G β€” World-model-mediated

  • Mechanism: Train policies inside a video world model that already abstracts embodiment.
  • Exemplars: DreamGen, Cosmos Policy, Ctrl-World.
  • Pros: βœ… Infinite synthetic data Β· βœ… shared visual prior across robots Β· βœ… new-robot deployment without real data.
  • Cons: ❌ Pixel world models not yet precise enough for contact-heavy skills Β· ❌ heavy compute Β· ❌ sim-to-real still required.
  • Pick when: You have no real data for the target embodiment.

6. Cross-axis comparison

Axis Best in class
Cheapest new-robot extension X-VLA (B) β€” train a tiny soft prompt
Most diverse morphology span CrossFormer (E) β€” arms + wheels + legs + wings
Most data-efficient transfer UniVLA-2505 (C) β€” 1/20 compute, 1/10 data vs OpenVLA
Strongest zero-shot result in 2026 Ο€0.7 (F) β€” UR5e laundry β‰ˆ expert teleop
Most principled action-space unification RDT-1B (A) Physically Interpretable UAS, UniAct (C)
Best for humanoid whole-body WholeBodyVLA (E), GR00T N1.6 (E+F)
Best for dex hands with no teleop DexUMI (D)
Best for "no real data, only video" Visual Imitation β†’ Humanoid (D+G), DreamGen (G)
Canonical falsifier of claims AnyBody benchmark β€” shows novel-morphology still fails

7. Research trends (2023 β†’ 2026)

  1. 2023 β€” "can we even share?" RT-2 showed web VLMs help robots. Cross-embodiment was aspirational; teams hand-crafted per-robot stacks.
  2. Late 2023 / 2024 β€” Group A (unified tokens) dominates. OXE + RT-1-X / RT-2-X demonstrate positive transfer across 22 robots. OpenVLA opens the recipe. Working assumption: "tokenize EE-pose deltas, pretrain on everything."
  3. 2024 β€” Group E (morphology-aware architecture) appears. Octo, CrossFormer, HPT factor obs/action spaces by embodiment. The reply to RT-X: "we need architecture, not just data."
  4. Early-mid 2025 β€” Group C (invariant latents) blooms. UniAct atomic codes, UniVLA latent actions, TraceVLA visual traces, X-Sim object motion, UniSkill skill embeddings. Thesis: "don't unify the action β€” unify something that doesn't depend on the body."
  5. Mid 2025 β€” Group D (human video / wearable) mainstreams. DexUMI, EgoDex, Visual Imitation β†’ Humanoid, ImMimic. Ο€0.5 co-trains on web + human data. "The best cross-embodiment dataset is the human body."
  6. Late 2025 β€” Group F (scale + prompt expansion) crowns production. Ο€0.5 β†’ Ο€0.6 (Nov 2025) β†’ Ο€0.7 (Apr 2026). Gemini Robotics 1.5. Thesis: "give the model enough data + rich prompts and it figures out the embodiment by itself." Ο€0.7 zero-shot UR5e is the headline.
  7. Late 2025 / 2026 β€” ICLR 2026 synthesizes B + C + G. X-VLA = clean scaling laws via soft prompts; XR-1 = shared codebook across vision+motion; WholeBodyVLA = humanoid integration; DreamGen / Cosmos Policy / Ctrl-World = world models inside training.

Direction of travel:

  • Mechanism consolidation. Soft-prompt (B) + invariant latent (C) + rich prompt (F) are converging β€” X-VLA's prompts and Ο€0.7's metadata are the same idea from different angles.
  • Data substrate unification. OXE + human video (EgoDex) + world-model rollouts (DreamGen / Cosmos) + RL traces (RECAP) is becoming the default pretraining mix.
  • Benchmarks catch up. AnyBody (2025) makes novel-morphology generalization falsifiable β€” and shows we're not there yet on extrapolation/composition.
  • Humanoid is in the tent. GR00T, WholeBodyVLA, Visual Imitation β†’ Humanoid bring whole-body morphology into the same VLA discussion as tabletop.

8. Open questions

Q1. Does scale alone solve cross-embodiment? Ο€0.7's UR5e result suggests scale + prompting is enough for robots that "look like" your training distribution. But AnyBody shows extrapolation to new link structures still fails. ICRA 2026 adds a third caveat from the data side: Galaxea / G0 cross-embodiment-pretrains on ~1,000 h OXE but finds the single-embodiment stage decisive, and few-shot tests favor embodiment-consistent pretraining β€” i.e., raw cross-embodiment scale is not automatically better than consistent, well-annotated data for a fixed target. Updated answer: scale for coverage, architecture for extrapolation, and data consistency/quality for the deployment embodiment β€” three levers, not one. The "give it enough data and it figures out the body itself" thesis (trend #6) holds for in-class morphologies but is not a substitute for target-embodiment data quality.

Q2. Sim-to-real vs. real-to-sim for cross-embodiment?

  • Sim-to-real (DreamGen, Cosmos Policy, Visual-Imit-Humanoid): best when task-completion is specifiable in sim and the embodiment gap is large (humanoid).
  • Real-to-sim-to-real (X-Sim): best when human demos exist and tasks are object-motion-reducible.
  • Real-only (Ο€0.7, Gemini 1.5): best when frontier-lab data budgets are available.
  • 2026 consensus: all three coexist; world-model-mediated sim (Group G) is closing the gap faster than expected.

Q3. Is there a universal action space that actually works?

  • EE-pose-delta tokenized (RT-X, Octo) β€” works for single-arm, breaks on dex and humanoid.
  • Physically Interpretable Unified Action Space (RDT-1B) β€” strictly better, still arm-centric.
  • Atomic action codes (UniAct) β€” promising at coarse-skill level, unclear at contact-rich level.
  • Latent action tokens (UniVLA-2505, XR-1 UVMC) β€” cleanest abstraction; tied to learned-feature quality.
  • Empirical answer in 2026: no single representation wins across all skills. Production systems (Ο€0.7) just condition on control-mode (joint vs EE) and let the flow-matching expert absorb the rest.

Q4. Can you transfer strategy, not just motion? Ο€0.7 folding UR5e shirts reportedly discovers new grasp strategies not seen in training β€” strong strategy-transfer signal. UniSkill and UniAct also argue for strategy transfer via skill/atomic codes. But:

  • Most cross-embodiment successes are within a class of strategies already in the data.
  • Truly novel strategy invention (e.g., a parallel-jaw gripper discovering a pinch it has never seen) remains rare.
  • The Ο€0.7 UR5e result is the strongest 2026 evidence of genuine strategy transfer; academic reproducibility is the key question for 2026–2027.

Q5. Are benchmarks keeping up? AnyBody (2025), RoboArena ∞, RoboCasa365, WorldGym all push this direction. Expect a shakeout in 2026 where Group F claims get tested on AnyBody-style extrapolation suites β€” and we see which ones survive.

Q6. Closed vs. open ecosystems. The strongest cross-embodiment systems (Ο€0.7, Gemini Robotics 1.5) are closed-weights. The strongest open alternatives (GR00T N1.6, OpenVLA, Octo, RDT-1B) trail on the hardest tasks. Whether open ecosystems close this gap in 2026 depends on who ships the next large multi-embodiment dataset β€” and whether it comes from academia (OXE v2?), industry (NVIDIA?), or a coalition.


9. Practical decision guide

If you're shipping a cross-embodiment VLA and wondering where to spend a quarter:

  1. OXE-style data on a few single-arm robots? β†’ Start with Octo / OpenVLA / RDT-1B as baselines (Cat A). Don't overthink it.
  2. Many robots, want cheapest per-new-robot cost? β†’ X-VLA-style soft prompts on a shared backbone (Cat B).
  3. Big human-robot gap or human-video-abundant task? β†’ Pick an invariant-latent method (Cat C β€” UniAct or X-Sim), or bridge via human demos (Cat D β€” DexUMI, EgoDex, ImMimic).
  4. Legs / wheels / wings in the mix? β†’ Go morphology-aware (Cat E β€” CrossFormer or GR00T N1).
  5. Humanoid whole-body target? β†’ WholeBodyVLA (Cat E) or GR00T N1.6 (Cat E+F) with paired loco-manipulation data.
  6. Have frontier-scale data budget? β†’ Follow the Ο€0.7 / Gemini 1.5 playbook (Cat F): prompt expansion + massive multi-robot mix + RL rollouts distilled back in.
  7. No real data for the target embodiment? β†’ DreamGen or Cosmos Policy rollouts (Cat G) as a bootstrap; fine-tune on the real target once you have any data.

Don't skip benchmarking on AnyBody-style extrapolation. If your system only works on robots that look like the training distribution, that's fine β€” but the deployment story differs a lot from "universal."


10. Links & wiki cross-refs


πŸ—“ State of the Field (updated Aug 2026)

Verdict: 2026 flipped the question β€” action-representation alignment is the precondition for cross-embodiment data to scale at all; hands joined the agenda with 80%+ zero-shot transfer.

πŸ“ˆ Trend

Qwen-RobotManip's controlled scaling ablation is the pivot: naive zero-padded spaces yield no scaling law; a canonical 80-dim representation scales; camera-frame delta EEF scales log-linearly and transfers zero-shot to unseen arms (RoboTwin-XE 23.9% vs 14.5% joint; UR5 5.6Γ— better). RSS 2026 extended to hands: DexGrasp-Zero (morphology-aligned graphs + fixed per-hand mapping, 85% zero-shot, +59.5% over SOTA) and One-Hand (canonical URDF + smooth morphology latent, 81.9%); LAP-style language-action pretraining and X-DiffVLA-class cross-embodied heads round out the cluster.

βš–οΈ Approaches & trade-offs

Interface Pros Cons Proven at
Canonical padded tensor + masks Simple, scales, one policy Slot semantics curated per embodiment Qwen suite, GR00T-class
Camera-frame delta EEF (+CaPE) Visual-action alignment; best zero-shot; creates the scaling law Calibrated cameras needed at train & test RobotManip
Morphology graphs / canonical URDFs Anatomy-grounded; no trainable retargeting Hands-only so far DexGrasp-Zero, One-Hand
Embodiment prompts + in-context history No architecture change Weak alone (+2–3 pp; history costs denoise steps) RobotManip ablations
Latent action/motion codes (ICML 2026) Embodiment-agnostic by construction Decoding fidelity per platform XR-1 (Oral: unified vision-motion codes, 12k+ rollouts / 6 embodiments); LAC-WM (+46.7% unseen-robot WM adaptation)
Data-side augmentation (ICML 2026) Works with any interface Synthetic-render gap OXE-AugE re-renders OXE across 9 embodiments (4.4M trajectories; +24–45% on unseen robot-gripper combos); RDT2 scales 10k+ h UMI data to zero-shot cross-embodiment

⚠️ Limitations & open problems

  • Joint-space zero-shot transfer still fails (<5% on unseen morphologies).
  • Morphology distance matters: Franka-class 7-DoF gaps resist even camera-frame transfer (5.9%); gripper↔dexterous-hand untried.
  • Calibration is a hidden deployment tax for the best interface; uncalibrated fallbacks under-evaluated.
  • Human↔robot is the hardest pair β€” see Review-Human-Video-Transfer.

← Back to Home

⚠️ **GitHub.com Fallback** ⚠️