Review Single Checkpoint Multi Robot - Heungwoo/research GitHub Wiki

In-Depth Review β€” Single-Checkpoint Multi-Robot Deployment (one frozen model, many robots at inference)

The property this page tracks: at inference, one unchanged checkpoint issues valid commands to two or more physically different robots β€” ideally without swapping weights or fine-tuning per robot. Distinct from Cross-Embodiment (training), which organizes papers by how heterogeneous data is mixed at training time. A model can be cross-embodiment-trained yet still need per-robot fine-tuning to deploy (e.g. Octo). This page asks only the deployment question. Companions: VLA Hybrid Architectures Β· World Models Β· Humanoid VLA.


1. The deployment question, decomposed

Training-time cross-embodiment (data mixing) and deploy-time single-checkpoint control are different properties. The deploy question has three parts:

  1. How is the target robot selected at inference? β€” the routing mechanism baked into the frozen checkpoint (unified action space Β· per-embodiment head Β· soft prompt Β· metadata/control-mode token Β· actions-as-language Β· embodiment-aware encoder).
  2. What is the zero-shot scope? β€” does one checkpoint control only the seen robots it trained on, or does it transfer zero-shot to an unseen robot?
  3. What ships per robot? β€” the honest asterisk: many "single checkpoint" systems still carry a per-robot prompt / readout head / calibration inside the checkpoint. Truly nothing-per-robot is the rare, hard case.

2. Three deployment tiers

  • Tier 1 β€” routed seen-robot generalist. One checkpoint controls all robots in its training mix; the robot is selected by a routing signal (head / token / action-space). Zero-shot to unseen robots: no.
  • Tier 2 β€” zero-shot to an unseen robot. One frozen checkpoint controls a robot it never trained on. This is the 2026 frontier (mostly recent, mostly unreplicated).
  • Tier 3 β€” one backbone, but not single-checkpoint deploy (contrast class). Cross-embodiment-trained, but a new robot needs a new head/stem/prompt fit from demos β€” so deployment is not a single frozen checkpoint. Included to sharpen the boundary.

3. Comparison

Tier 2 β€” one frozen checkpoint β†’ an unseen robot (the frontier):

Model Robots Inference routing Evidence Source
LAP-3B novel robots + tasks actions as natural language β€” no tokenizer, no per-embodiment head "first VLA to achieve substantial zero-shot transfer to previously unseen robot embodiments"; >50% avg zero-shot, ~2Γ— prior VLAs 2602.10556 Β· RSS-2026-LAP
Gemini Robotics 1.5 ALOHA · bi-arm Franka · Apollo humanoid Motion-Transfer pretraining objective + prompt zero-shot skill transfer across all three without robot-specific post-training (ALOHA→Apollo is a large morphology jump) 2510.03342 (closed)
Green-VLA humanoids Β· mobile manipulators Β· fixed-base arms unified embodiment-aware action interface; 5-stage curriculum (L0 VLM β†’ L1 grounding β†’ R0 multi-embodiment β†’ R1 adapt β†’ R2 RL) single policy across morphologies; zero-shot to new embodiments; Simpler-BRIDGE / CALVIN + real 2602.00919
[DreamZero](/Heungwoo/research/wiki/Review-DreamZero) (WAM) diverse video-diffusion world-action backbone "World Action Models are Zero-shot Policies" β€” boosts task and embodiment generalization 2602.15922
[DYNA-2](/Heungwoo/research/wiki/Review-Dyna2) (WAM, vendor) fixed arm Β· humanoid Β· 5-finger dex hand joint next-frame+action WAM vendor-claimed zero-shot production-level cross-embodiment (unverified) dyna.co
[Contact-Anchored Policies](/Heungwoo/research/wiki/RSS-2026-Contact-Anchored-Policies) 3 skills, novel env + embodiment contact-point conditioning (not language) out-of-box on unseen embodiments; +56% over big VLAs zero-shot on 23 h data RSS 2026
[One Hand to Rule Them All](/Heungwoo/research/wiki/RSS-2026-One-Hand) unseen dexterous hands canonical-URDF + learnable morphology latent 81.9% zero-shot on an unseen 3-finger hand RSS 2026

Tier 1 β€” one checkpoint β†’ all its seen robots (routed):

Model Robots Inference routing Note Source
Open X-Embodiment / RT-X 22 robots unified EE-pose text tokens (one shared space, no per-robot head) the original one-checkpoint-many-arms result RSS 2024
CrossFormer 20 embodiments incl. quadcopter, quadruped per-embodiment readout heads (all in the checkpoint) most morphology-diverse single net CoRL 2024
RDT-1B bimanual arms unified (physically-interpretable) action space + diffusion 1.2B bimanual foundation ICLR 2025
UniAct 28 embodiments universal atomic-action codes + per-embodiment decoder a 0.5B model reportedly beats 14Γ— larger baselines CVPR 2025
[GR00T N1](/Heungwoo/research/wiki/Review-GR00T-Series) humanoids + arms embodiment-aware state/action encoder open humanoid FM; needs per-robot calibration data 2025
[Ο€0.5](/Heungwoo/research/wiki/CoRL-2025-pi05) β†’ [Ο€0.7](/Heungwoo/research/wiki/PI-pi07) single/bi-arm, mobile, UR5e metadata + control-mode token (joint vs EE) in the prompt zero-shot UR5e shirt-fold 85.6% β‰ˆ expert teleop β€” strongest "scale+prompt solves in-class morphology" datapoint PI 2026
[Motus](/Heungwoo/research/wiki/Review-MOTUS) multi-robot scheduled MoT (Phase-3 specializes to the target robot) unified checkpoint, but target-robot phase is a fit 2512.13030

Tier 3 β€” contrast: one backbone, but a new robot needs a new fit (not single-checkpoint deploy):

Model Per-robot cost Source
Octo fine-tune a new head (hours) CoRL 2024
HPT train a new stem per embodiment NeurIPS 2024
[X-VLA](/Heungwoo/research/wiki/ICLR-2026-X-VLA) fit a new soft prompt from that robot's demos (backbone frozen) ICLR 2026

(An orthogonal route to one checkpoint: MergeVLA (2511.18810) merges specialist checkpoints into a single generalist agent β€” one checkpoint by construction, not by co-training.)


4. The routing mechanisms β€” how a frozen checkpoint targets a robot

Ordered roughly from "most per-robot machinery" to "none":

  1. Per-embodiment readout head / decoder β€” CrossFormer, UniAct, Octo. Each robot has its own output head inside the checkpoint; a selector routes to it.
  2. Embodiment-aware encoder β€” GR00T embeds robot-specific state/action; still often needs calibration data.
  3. Soft prompt per embodiment β€” X-VLA, HPT stems. Tiny per-robot parameters, but they must be fit (so not zero-shot).
  4. Metadata / control-mode token β€” Ο€0.7 (joint-vs-EE token + episode metadata). Routing is text, so a new in-class robot can be addressed zero-shot.
  5. Unified action space β€” RT-X (EE-pose deltas), RDT-1B (unified schema). No per-robot component, but the space is a lowest-common-denominator (weak for dex / legged).
  6. Actions as natural language β€” LAP. No tokenizer, no head, no prompt-fit β€” the VLM emits actions as text, so nothing is per-robot. This is why LAP is the cleanest zero-shot-to-unseen claim.
  7. World model as the policy β€” DreamZero, DYNA-2. The shared video/dynamics prior generalizes across bodies; the action head is thin or reactive.

5. Honest limits

  • Unseen-morphology extrapolation is still largely unsolved in controlled tests. The AnyBody benchmark (2505.14986) shows in-distribution cross-embodiment works but genuinely novel morphology / composition fails β€” the canonical falsifier every Tier-2 claim must face.
  • Tier-2 is new and mostly unreplicated. LAP-3B, Green-VLA (Jan–Feb 2026 preprints), Gemini Robotics 1.5 (closed), DYNA-2 (vendor) β€” none has independent third-party cross-robot replication yet. Treat "zero-shot to unseen" as promising-but-provisional.
  • "Single checkpoint" β‰  "nothing per robot." Most Tier-1 systems still ship a per-robot head/prompt/encoder inside the checkpoint. Only actions-as-language (LAP) and unified-action-space (RT-X/RDT) carry truly no per-robot component β€” at the cost of expressiveness.
  • The action-space tax. Unified lowest-common-denominator spaces (EE-pose) deploy everywhere but under-serve dexterous hands and legged/whole-body control β€” where per-embodiment heads or WAM priors still win.

6. Design guide β€” I need one checkpoint for N robots

  1. All robots known at build time, similar action spaces? β†’ unified action space (RT-X/RDT) or routed heads (CrossFormer/UniAct). Simplest, robust, seen-only.
  2. New robots will appear, and they're in-class (another arm/bimanual)? β†’ metadata/prompt conditioning (Ο€0.7 style). Address new-in-class robots by text, no new parameters.
  3. You need zero-shot to a genuinely unseen robot, no demos? β†’ actions-as-language (LAP) or a WAM policy (DreamZero) β€” the only routes with nothing per-robot. Accept lower precision and provisional evidence.
  4. Very different morphologies (arm ↔ humanoid ↔ mobile)? β†’ unified embodiment-aware interface + staged curriculum (Green-VLA) or Motion-Transfer pretraining (Gemini Robotics 1.5).
  5. Have strong per-robot specialists already? β†’ merge them (MergeVLA) instead of retraining a joint model.
  6. Always benchmark against AnyBody before claiming universality.

7. Links

← Back to Reviews Β· Home