Review Single Checkpoint Multi Robot - Heungwoo/research GitHub Wiki
In-Depth Review β Single-Checkpoint Multi-Robot Deployment (one frozen model, many robots at inference)
The property this page tracks: at inference, one unchanged checkpoint issues valid commands to two or more physically different robots β ideally without swapping weights or fine-tuning per robot. Distinct from Cross-Embodiment (training), which organizes papers by how heterogeneous data is mixed at training time. A model can be cross-embodiment-trained yet still need per-robot fine-tuning to deploy (e.g. Octo). This page asks only the deployment question. Companions: VLA Hybrid Architectures Β· World Models Β· Humanoid VLA.
1. The deployment question, decomposed
Training-time cross-embodiment (data mixing) and deploy-time single-checkpoint control are different properties. The deploy question has three parts:
- How is the target robot selected at inference? β the routing mechanism baked into the frozen checkpoint (unified action space Β· per-embodiment head Β· soft prompt Β· metadata/control-mode token Β· actions-as-language Β· embodiment-aware encoder).
- What is the zero-shot scope? β does one checkpoint control only the seen robots it trained on, or does it transfer zero-shot to an unseen robot?
- What ships per robot? β the honest asterisk: many "single checkpoint" systems still carry a per-robot prompt / readout head / calibration inside the checkpoint. Truly nothing-per-robot is the rare, hard case.
2. Three deployment tiers
- Tier 1 β routed seen-robot generalist. One checkpoint controls all robots in its training mix; the robot is selected by a routing signal (head / token / action-space). Zero-shot to unseen robots: no.
- Tier 2 β zero-shot to an unseen robot. One frozen checkpoint controls a robot it never trained on. This is the 2026 frontier (mostly recent, mostly unreplicated).
- Tier 3 β one backbone, but not single-checkpoint deploy (contrast class). Cross-embodiment-trained, but a new robot needs a new head/stem/prompt fit from demos β so deployment is not a single frozen checkpoint. Included to sharpen the boundary.
3. Comparison
Tier 2 β one frozen checkpoint β an unseen robot (the frontier):
| Model | Robots | Inference routing | Evidence | Source |
|---|---|---|---|---|
| LAP-3B | novel robots + tasks | actions as natural language β no tokenizer, no per-embodiment head | "first VLA to achieve substantial zero-shot transfer to previously unseen robot embodiments"; >50% avg zero-shot, ~2Γ prior VLAs | 2602.10556 Β· RSS-2026-LAP |
| Gemini Robotics 1.5 | ALOHA Β· bi-arm Franka Β· Apollo humanoid | Motion-Transfer pretraining objective + prompt | zero-shot skill transfer across all three without robot-specific post-training (ALOHAβApollo is a large morphology jump) | 2510.03342 (closed) |
| Green-VLA | humanoids Β· mobile manipulators Β· fixed-base arms | unified embodiment-aware action interface; 5-stage curriculum (L0 VLM β L1 grounding β R0 multi-embodiment β R1 adapt β R2 RL) | single policy across morphologies; zero-shot to new embodiments; Simpler-BRIDGE / CALVIN + real | 2602.00919 |
| [DreamZero](/Heungwoo/research/wiki/Review-DreamZero) (WAM) | diverse | video-diffusion world-action backbone | "World Action Models are Zero-shot Policies" β boosts task and embodiment generalization | 2602.15922 |
| [DYNA-2](/Heungwoo/research/wiki/Review-Dyna2) (WAM, vendor) | fixed arm Β· humanoid Β· 5-finger dex hand | joint next-frame+action WAM | vendor-claimed zero-shot production-level cross-embodiment (unverified) | dyna.co |
| [Contact-Anchored Policies](/Heungwoo/research/wiki/RSS-2026-Contact-Anchored-Policies) | 3 skills, novel env + embodiment | contact-point conditioning (not language) | out-of-box on unseen embodiments; +56% over big VLAs zero-shot on 23 h data | RSS 2026 |
| [One Hand to Rule Them All](/Heungwoo/research/wiki/RSS-2026-One-Hand) | unseen dexterous hands | canonical-URDF + learnable morphology latent | 81.9% zero-shot on an unseen 3-finger hand | RSS 2026 |
Tier 1 β one checkpoint β all its seen robots (routed):
| Model | Robots | Inference routing | Note | Source |
|---|---|---|---|---|
| Open X-Embodiment / RT-X | 22 robots | unified EE-pose text tokens (one shared space, no per-robot head) | the original one-checkpoint-many-arms result | RSS 2024 |
| CrossFormer | 20 embodiments incl. quadcopter, quadruped | per-embodiment readout heads (all in the checkpoint) | most morphology-diverse single net | CoRL 2024 |
| RDT-1B | bimanual arms | unified (physically-interpretable) action space + diffusion | 1.2B bimanual foundation | ICLR 2025 |
| UniAct | 28 embodiments | universal atomic-action codes + per-embodiment decoder | a 0.5B model reportedly beats 14Γ larger baselines | CVPR 2025 |
| [GR00T N1](/Heungwoo/research/wiki/Review-GR00T-Series) | humanoids + arms | embodiment-aware state/action encoder | open humanoid FM; needs per-robot calibration data | 2025 |
| [Ο0.5](/Heungwoo/research/wiki/CoRL-2025-pi05) β [Ο0.7](/Heungwoo/research/wiki/PI-pi07) | single/bi-arm, mobile, UR5e | metadata + control-mode token (joint vs EE) in the prompt | zero-shot UR5e shirt-fold 85.6% β expert teleop β strongest "scale+prompt solves in-class morphology" datapoint | PI 2026 |
| [Motus](/Heungwoo/research/wiki/Review-MOTUS) | multi-robot | scheduled MoT (Phase-3 specializes to the target robot) | unified checkpoint, but target-robot phase is a fit | 2512.13030 |
Tier 3 β contrast: one backbone, but a new robot needs a new fit (not single-checkpoint deploy):
| Model | Per-robot cost | Source |
|---|---|---|
| Octo | fine-tune a new head (hours) | CoRL 2024 |
| HPT | train a new stem per embodiment | NeurIPS 2024 |
| [X-VLA](/Heungwoo/research/wiki/ICLR-2026-X-VLA) | fit a new soft prompt from that robot's demos (backbone frozen) | ICLR 2026 |
(An orthogonal route to one checkpoint: MergeVLA (2511.18810) merges specialist checkpoints into a single generalist agent β one checkpoint by construction, not by co-training.)
4. The routing mechanisms β how a frozen checkpoint targets a robot
Ordered roughly from "most per-robot machinery" to "none":
- Per-embodiment readout head / decoder β CrossFormer, UniAct, Octo. Each robot has its own output head inside the checkpoint; a selector routes to it.
- Embodiment-aware encoder β GR00T embeds robot-specific state/action; still often needs calibration data.
- Soft prompt per embodiment β X-VLA, HPT stems. Tiny per-robot parameters, but they must be fit (so not zero-shot).
- Metadata / control-mode token β Ο0.7 (joint-vs-EE token + episode metadata). Routing is text, so a new in-class robot can be addressed zero-shot.
- Unified action space β RT-X (EE-pose deltas), RDT-1B (unified schema). No per-robot component, but the space is a lowest-common-denominator (weak for dex / legged).
- Actions as natural language β LAP. No tokenizer, no head, no prompt-fit β the VLM emits actions as text, so nothing is per-robot. This is why LAP is the cleanest zero-shot-to-unseen claim.
- World model as the policy β DreamZero, DYNA-2. The shared video/dynamics prior generalizes across bodies; the action head is thin or reactive.
5. Honest limits
- Unseen-morphology extrapolation is still largely unsolved in controlled tests. The AnyBody benchmark (2505.14986) shows in-distribution cross-embodiment works but genuinely novel morphology / composition fails β the canonical falsifier every Tier-2 claim must face.
- Tier-2 is new and mostly unreplicated. LAP-3B, Green-VLA (JanβFeb 2026 preprints), Gemini Robotics 1.5 (closed), DYNA-2 (vendor) β none has independent third-party cross-robot replication yet. Treat "zero-shot to unseen" as promising-but-provisional.
- "Single checkpoint" β "nothing per robot." Most Tier-1 systems still ship a per-robot head/prompt/encoder inside the checkpoint. Only actions-as-language (LAP) and unified-action-space (RT-X/RDT) carry truly no per-robot component β at the cost of expressiveness.
- The action-space tax. Unified lowest-common-denominator spaces (EE-pose) deploy everywhere but under-serve dexterous hands and legged/whole-body control β where per-embodiment heads or WAM priors still win.
6. Design guide β I need one checkpoint for N robots
- All robots known at build time, similar action spaces? β unified action space (RT-X/RDT) or routed heads (CrossFormer/UniAct). Simplest, robust, seen-only.
- New robots will appear, and they're in-class (another arm/bimanual)? β metadata/prompt conditioning (Ο0.7 style). Address new-in-class robots by text, no new parameters.
- You need zero-shot to a genuinely unseen robot, no demos? β actions-as-language (LAP) or a WAM policy (DreamZero) β the only routes with nothing per-robot. Accept lower precision and provisional evidence.
- Very different morphologies (arm β humanoid β mobile)? β unified embodiment-aware interface + staged curriculum (Green-VLA) or Motion-Transfer pretraining (Gemini Robotics 1.5).
- Have strong per-robot specialists already? β merge them (MergeVLA) instead of retraining a joint model.
- Always benchmark against AnyBody before claiming universality.
7. Links
- Tier-2 (zero-shot-to-unseen): LAP (2602.10556) Β· Green-VLA (2602.00919) Β· Gemini Robotics 1.5 (2510.03342) Β· DreamZero Β· DYNA-2 Β· Contact-Anchored Policies Β· One-Hand
- Tier-1 (routed seen-robot): RT-X Β· CrossFormer Β· RDT-1B Β· UniAct Β· GR00T N1 Β· Ο0.7 Β· Motus
- Contrast / adjacent: X-VLA Β· XR-1 Β· UniVLA Β· WholeBodyVLA
- The training-side companion: Cross-Embodiment (training) Β· benchmark: AnyBody (2505.14986)