Review LBM Cotraining - Heungwoo/research GitHub Wiki
Paper: A Systematic Study of Data Modalities and Strategies for Co-training Large Behavior Models for Robot Manipulation Lead author: Fanqi Lin Β· Senior author / corresp.: Jose Barreiros ([email protected]) Co-authors include: Kushal Arora, Jean Mercat, Haruki Nishimura, Paarth Shah, Chen Xu, β¦ Affiliation: Toyota Research Institute (TRI) arXiv: 2602.01067 Β· Feb 1, 2026 Code / data: internal TRI; no public release announced as of Feb 2026
This page sits in the VLA architectures review as a training-recipe contribution and pairs with Knowledge Insulation (the canonical "how to protect the VLM during action training" result) as the what data should you co-train on companion.
Toyota Research Institute trains 89 LBM (Large Behavior Model) policies β a flow-matching VLA built on PaliGemma2-3B β under a controlled 5Γ3 grid of data modality Γ training strategy, and runs 58,000 simulation rollouts + 2,835 real-robot rollouts to measure each combination.
Three findings overturn folk wisdom that had been propagating through 2025:
- Web vision-language data and cross-embodiment robot data dominate β they account for almost all of the +36 pp simulation gain and +45 pp real-world language-following gain.
- Discrete action tokens (FAST, VQ-VAE) and latent video tokens do not help at LBM scale β possibly hurt β despite being widely adopted in 2025 VLAs.
- Explicit chain-of-thought conditioning at inference is harmful for direct perception-to-action tasks; the gain people attributed to CoT is from training-time supervision, not inference-time conditioning. This empirically corroborates from a different angle the Hybrid Training thesis.
Two further findings refine the recipe:
- Cross-embodiment data is best used in phase 1 only β keeping it in phase 2 hurts.
- Robot-only training erodes the VLM's MMBench/GQA scores; balanced co-training restores them β i.e., VL co-training is partly a regularization-against-forgetting story, not just a generalization-gain story.
By Q1 2026 the community had three loud co-training claims fighting for default-recipe status:
- PI / Knowledge Insulation (NeurIPS 2025): the gradient from the action expert must not back-prop into the VLM, or you lose the VLM's pretrained knowledge.
- OpenVLA / FAST / OmniSAT: discrete action tokenization is the way to share parameters between vision-language and action.
- ECoT family: chain-of-thought traces (text or latent) attached to actions improve generalization.
The TRI paper is the first controlled comparison that grades all three with the same backbone, the same robot platform, and the same evaluation suite. The verdict for each:
| 2025 claim | TRI 2026 verdict |
|---|---|
| Co-training with VL data helps | β confirmed β largest single-modality gain |
| Discrete action tokenization helps | β no significant benefit at LBM scale |
| Cross-embodiment data helps | β β but only in phase 1 |
| Chain-of-thought conditioning at inference helps | β does not improve and often degrades for direct manipulation |
| Latent video tokens help | β marginal/null at scale; useful only in low-data regime |
| Robot-only training is fine | β erodes the VLM's visiolinguistic ability |
This is the empirical companion paper that the field had been waiting for since the Sergey Levine et al. What Matters in Learning from Large-Scale Datasets (ICRA 2025) line β same spirit, but at LBM scale and including the language/CoT axes that the older paper did not cover.
flowchart LR
IMG[Multi-camera observation] --> VLM
LANG[Language instruction] --> VLM
STATE[Proprio state] --> VLM
VLM[PaliGemma2-3B-PT VLM<br/>trained jointly with action expert]
VLM -- single observation token,<br/>extracted from last 4 layers --> COND[adaLN conditioning]
COND --> AFT[ActionFT β 8-layer flow-matching transformer]
AFT --> ACT[Action chunk: 16 steps<br/>relative EE pose + gripper width]
VLM -.discrete-token CE losses.-> AUX[Auxiliary discrete heads<br/>FAST tokens / VQ-VAE / language captions]
classDef vlm fill:#bbdefb,stroke:#1565c0,color:#000
classDef act fill:#c8e6c9,stroke:#2e7d32,color:#000
classDef aux fill:#fff9c4,stroke:#f57f17,color:#000
class VLM,COND vlm
class AFT,ACT act
class AUX aux
- Backbone: PaliGemma2-3B-PT, trained jointly with the action expert (not frozen) β the VLM "can optionally be trained to generate text or discrete action tokens." This is essential to the paper's catastrophic-forgetting result: only because the VLM weights move can robot-only training erode its MMBench/GQA scores and VL co-training restore them.
- Action expert: ActionFT, an 8-layer flow-matching transformer, conditioned via adaLN on a single observation token built from the last 4 layers of the VLM.
- Action representation: continuous flow-matching for the primary head; auxiliary discrete heads added under specific co-training settings.
- Action horizon: 16-step chunks.
Loss = flow-matching MSE on continuous actions + cross-entropy on auxiliary token / language streams when present, weighted by computed masks.
Note (correction): an earlier version of this page described the VLM as "frozen during action training." The source describes a jointly-trained end-to-end system; the VLM is not frozen.
This architecture is deliberately close to the public PaliGemma-3B + flow-matching recipe so the conclusions transfer to other LBM-class systems built on the same template (Ο0.5 / Ο0.6 / Steerable-Ο0.5, GR00T-N1.x family modulo backbone size).
flowchart TB
subgraph M1[Modality 1 β Standard VL data Β· 50M samples]
direction LR
RP[RoboPoint<br/>1.3M images Β· 8.2M QA]
RS[RefSpatial<br/>2.5M images Β· 20M QA]
end
subgraph M2[Modality 2 β Dense robot annotations Β· 523h]
direction LR
SC[Scripted heuristics<br/>EE-motion descriptions]
GC[GPT-5 captions<br/>rich semantic descriptions]
end
subgraph M3[Modality 3 β Cross-embodiment robot Β· 1,150h]
OXE[OXE-Ramen<br/>12 robots Β· 924 tasks Β· 466k demos]
end
subgraph M4[Modality 4 β Human videos Β· 2,271h]
direction LR
LAM[Latent Action Model<br/>32 codebook Β· 8 tokens/chunk]
HC[GPT-5 captions of motion]
end
subgraph M5[Modality 5 β Discrete robot tokens Β· 523h]
direction LR
FAST[FAST tokens<br/>~42 / chunk Β· vocab 2,048]
VQ[VQ-VAE tokens<br/>8 / chunk Β· codebook 32]
end
classDef good fill:#c8e6c9,stroke:#2e7d32,color:#000
classDef bad fill:#ffcdd2,stroke:#c62828,color:#000
class M1,M3 good
class M2 good
class M5 bad
Color coding: green = the paper's "effective" verdict, red = "no significant benefit", mixed for M2/M4 (caption variants effective, token variants not).
- Sources: RoboPoint (1.3M images, 8.2M QA pairs β robotics-relevant pointing and spatial QA) and RefSpatial (2.5M images, 20M QA β refer-and-locate spatial reasoning). Combined β 50M VL samples.
- Loss: standard CE on the VLM language head.
- Verdict: β largest single-modality contribution to both unseen-task generalization (sim) and language following (real). Cheap to scale because no robot data collection is required.
- Source: TRI-Ramen (~523 h of TRI's in-house manipulation demos β 403 tasks, 53,411 demonstrations on the dual-Franka platform).
-
Two caption variants:
- Scripted heuristics β derive end-effector motion descriptions from action stream ("gripper moves left and down 5 cm, closes").
- GPT-5 captions β vision-language model emits rich semantic captions for short trajectory segments at ~1β2 second intervals.
- Verdict: β GPT-5 captions clearly beat scripted heuristics. The semantic richness β object names, contact events, intent β is what matters. Scripted heuristics give the action a redundant motion description that the VLM does not need.
- Source: OXE-Ramen β a TRI-curated subset of Open X-Embodiment covering 12 robot setups, 924 tasks, 466k demos, totaling 1,150 h.
- Used as: native continuous actions (not retokenized).
- Verdict: β effective, but best used in phase 1 only. Keeping it in phase 2 hurts β the morphology gap between the OXE robots and the target dual-Franka platform creates conflicting gradients during specialization.
- Sources (2,271 h after filtering): Ego4D, EgoDex, Something-Something V2, Epic Kitchen, HoloAssist.
-
Two representations tested:
- Latent action tokens β Latent Action Model (LAM) trained to compress short clips into 8 tokens per chunk over a codebook of 32; supervised via CE.
- GPT-5 motion captions β VLM-generated language descriptions of the human motion.
- Verdict: β caption variant works. β latent-action variant only helps in the low-data regime; gains evaporate at scale. The paper's framing: human videos contribute visual diversity through language, not through action-token transfer.
-
Source: TRI-Ramen retokenized two ways:
- FAST β DCT-based, β 42 tokens per 16-step chunk over a 2,048 vocab.
- VQ-VAE β 8 tokens per chunk over a 32 codebook.
- Used as: auxiliary CE prediction targets alongside the primary flow-matching head.
- Verdict: β no significant benefit; FAST sometimes hurts unseen-task performance. This is one of the paper's most important findings β Knowledge Insulation NeurIPS-2025-Knowledge-Insulation uses FAST tokens to give the VLM a CE signal, but TRI's controlled study shows the signal itself is not what matters; KI's gain must come from the gradient routing, not from the token target.
flowchart LR
subgraph SP[Single-phase joint]
SP1[Mix target + co-training<br/>train one phase end-to-end]
end
subgraph TP1[Two-phase Β· 1st-phase only]
TP1a[Phase 1: co-training data only]
TP1b[Phase 2: target robot continuous actions only]
TP1a --> TP1b
end
subgraph TP2[Two-phase Β· full co-training]
TP2a[Phase 1: co-training data only]
TP2b[Phase 2: target + co-training jointly]
TP2a --> TP2b
end
classDef ph fill:#fff9c4,stroke:#f57f17,color:#000
class SP1,TP1a,TP1b,TP2a,TP2b ph
The strategy that wins depends on the modality:
| Modality | Best strategy |
|---|---|
| Standard VL | Two-phase full (keep VL in phase 2) |
| Robot annotations (GPT-5) | Two-phase, 1st-phase only |
| Cross-embodiment robot | Two-phase, 1st-phase only |
| Human-video captions | Two-phase full (keep in phase 2) |
| FAST / VQ-VAE / latent video tokens | none of the strategies recover usable gain |
Pattern: language-bearing modalities want to be co-trained throughout (phase 1 + phase 2). Action-stream modalities from non-target morphologies want to be dropped after phase 1 so the policy can specialize.
This pattern is consistent with how the Ο-series uses Knowledge Insulation: the VLM keeps getting language-side gradients (CE on tokens / captions) throughout training, while the action expert specializes β TRI's empirical result is the data-side counterpart of KI's gradient-routing result.
- Tasks: 13 seen + 8 unseen (semantic / multi-step / compositional generalization).
- Conditions: nominal + four distribution shifts β lighting, textures, camera viewpoint, backgrounds.
- Rollouts: 50 per task per condition β 58,000+ total.
flowchart LR
subgraph LF[Language-following Β· 45 rollouts/condition]
LF1[Seen objects from training]
LF2[Instruction generalization<br/>paraphrases Β· category Β· attribute]
LF3[Unseen objects Β· 52 novel]
end
subgraph LH[Long-horizon dexterous Β· 30 rollouts each Β· ~93 s Β· 13 steps avg]
LH1[PackItemsIntoStringBag<br/>fine-grained bottle capping]
LH2[PourIngredientsIntoSoup<br/>scooping with spatula]
LH3[StoreCleanDishes<br/>transparent objects]
end
- Total real-robot rollouts: 2,835.
In addition to manipulation, the LBM's PaliGemma2 backbone is evaluated on MMBench, GQA, and other standard VL benchmarks before/after each training regime β to detect catastrophic forgetting of visiolinguistic capability.
| Setting | Unseen tasks |
|---|---|
| Baseline (target-only training) | β 36.2% |
| + Standard VL | ββ |
| + GPT-5 robot captions | β |
| + Cross-embodiment (phase 1 only) | β |
| + Human-video captions | β |
| All effective modalities stacked | 72.6% |
β +36.4 percentage points from the right co-training mix; gains stack additively rather than saturating.
| Setting | Task completion (avg over 3 conditions) |
|---|---|
| Baseline | β 24.1% |
| All effective modalities | 69.4% |
β +45.3 pp. The unseen-objects axis is the largest single beneficiary β VL data + human-video captions push the model to recognize objects it never saw on the target robot.
After pretraining, fine-tune on 200 demonstrations of three unseen long-horizon tasks (bottle-capping in bag, scooping into soup, storing transparent dishes):
| Setting | Completion |
|---|---|
| Single-task untrained baseline | 47.3% |
| Pretrained baseline (no co-training) β fine-tune | 67.4% |
| Full co-trained LBM β fine-tune | 90.2% |
The +22.8 pp delta over the same fine-tune procedure on the no-co-training baseline is the cleanest demonstration in the paper that co-training learns transferable representations, not just multi-task amortization.
Robot-only training drops PaliGemma2 MMBench / GQA scores measurably vs. the original VLM weights. Adding standard VL co-training restores them. This frames part of the gain as preventing catastrophic forgetting rather than learning new capability.
This is the section that will be most cited because it directly contradicts the prevailing 2024-2025 ECoT narrative.
Three CoT trace types are tested:
- Scripted captions (Modality 2a)
- VLM-generated captions (Modality 2b / Modality 4b)
- Latent action tokens (Modality 4a)
For each, three CoT strategies:
| # | Train | Inference |
|---|---|---|
| 1 | 50% with CoT, 50% without | Without CoT (skip generation) |
| 2 | 50% with CoT, 50% without | With CoT β generate trace then condition |
| 3 | Always with CoT | Always with CoT |
All three CoT inference strategies either match or degrade the implicit two-phase co-training baseline.
The paper's diagnosis:
- The eval tasks are direct perception-to-action (pick-and-place, simple long-horizon scoops). They do not require multi-step planning.
- VLM-generated captions hallucinate at non-trivial rates; latent action tokens contain reconstruction noise.
- When the action prediction is conditioned on the noisy generated CoT, errors propagate: a hallucinated subgoal directly distorts the action chunk that follows.
- When CoT is only used as an auxiliary CE training target and skipped at inference (option 1), the model still doesn't gain β the implicit co-training already captures whatever benefit was available.
This is the paper's most direct dialogue with the broader 2026 reasoning-augmented VLA literature:
- Hybrid Training had argued: the ECoT generalization gain is from training-time supervision; drop the reasoning at inference. TRI's controlled experiment empirically validates that thesis for the LBM-scale setting.
- Actions as Language had argued: subtask relabeling is the lever. TRI confirms: language captions of robot/human trajectories are the most valuable form of co-training, but only as CE training targets, not as inference conditioning.
- Embodied-R1 and dVLA are the two papers that do make CoT-at-inference work β they pay either via RL (E-R1) or via parallel decoding (dVLA). TRI's paper sharpens their justification: vanilla CoT-at-inference doesn't pay off; you need a specific architectural or training reason to use it.
This is the second most important finding for the field's recipes.
The 2024β2025 consensus had been that discrete action tokens (FAST, BPE, VQ-VAE) are the way to:
- Share parameters with VLM language heads (one CE loss).
- Make actions look like language so that a frozen LM can be co-trained.
- Compress high-frequency action streams.
Knowledge Insulation β the foundation of Ο0.6/Ο0.7 β uses FAST tokens specifically as the CE training target for the VLM half, while the continuous flow-matching head remains the action head.
TRI's controlled experiment shows that, with this exact architecture (PaliGemma2 + flow-matching + auxiliary FAST CE head), the FAST CE loss adds no measurable benefit and sometimes hurts unseen-task performance.
Two ways to reconcile:
- The KI gain comes from gradient routing, not from the token target. That is, what matters is that the VLM's gradient is insulated from the continuous head β the auxiliary CE target was a vehicle for that, not the cause of the gain.
- The benefit of FAST tokens depends on the eval suite. Possible β TRI's eval is dual-Franka manipulation; PI's flagship evals include very different domains (mobile manipulation, laundry folding). The TRI paper does not claim a universal verdict, only the LBM-scale verdict on its specific eval suite.
Either way, this is now an open empirical question the field needs to settle in 2026 β not a closed-recipe assumption.
Distilled from Β§III of the paper:
- Always include large-scale standard VL data (RoboPoint + RefSpatial or equivalent). It gives the largest single-modality gain and prevents VLM forgetting. Cheap; no robot data collection.
- Always include GPT-5-generated captions over your robot demos. Better than scripted heuristics. Keep them as a CE target throughout phase 2.
- Include cross-embodiment robot data (OXE) only in phase 1. Drop it for specialization in phase 2.
- Include human-video captions (not latent action tokens) and keep them throughout both phases.
- Skip discrete action tokenization (FAST / VQ-VAE) as an auxiliary head unless you have a specific reason (low-data regime, KI-style gradient routing argument).
- Skip explicit chain-of-thought conditioning at inference. If you want CoT's training-time benefit, use it as a CE auxiliary loss only and predict actions without generating CoT at inference.
- Two-phase training is the default. Single-phase joint training works too but does not match the two-phase ceiling.
This section consolidates everything needed to reproduce the Final Model ("+Cross-Embodiment-Robot-Data" row in Table S1) β the recipe that achieved the best generalization in the paper. Numbers are pulled from Appendix 2 (Training Details), Table S1 (per-phase data ratios), Table S2 (shared hyperparameters), and Table S3 (human-video data composition).
Note on "two-phase" vs "three-phase": The paper's main text (Fig. 9) reports that the best single-modality policies are either two-phase full co-training (standard VL, human-video annotations) or two-phase 1st-phase-only (scripted/VLM robot annotations, cross-embodiment data). That statement is about isolated modalities. The cumulative Final Model (Β§III-C, "+Cross-Embodiment-Robot-Data") is a separate model that stacks all effective modalities, and in Table S1 it has populated 1st-/2nd-/3rd-phase columns β i.e. it is genuinely three-phase. The two are not in conflict: "two-phase" is the per-modality finding; "three-phase" is the stacked-recipe schedule. (Audit 2026-06: all numbers in this section were cross-checked against the arXiv HTML full text incl. appendix and confirmed β see report.)
The best-performing policy is a three-phase schedule that mixes five modalities at carefully tuned ratios:
| Modality | Volume | Role |
|---|---|---|
| TRI-Ramen (target dual-Franka teleop) | 523 h | Continuous flow-matching anchor β the "golden" target distribution |
| OXE-Ramen (cross-embodiment) | 1,150 h, 466k demos, 12 robots, 924 tasks | Continuous flow-matching, phase-1+2 only (dropped in phase 3) |
| Standard VL (RoboPoint 8.2M QA + RefSpatial 20M QA) | β 50M VL samples | CE on language head; throughout all phases |
| GPT-5 language annotations of TRI-Ramen | β 523 h Γ 1β2 captions/s | CE on language head; throughout all phases |
| GPT-5 language annotations of human videos (Ego4D 774.5 h, EgoDex 744.4 h + reversed 455.7 h, Sth-Sth-V2 155.8 h, Epic Kitchen 60.4 h, HoloAssist 80.8 h β 2,271.6 h) | 9.0M annotations across 2,271.6 h | CE on language head; throughout all phases |
Excluded by ablation: FAST tokens, VQ-VAE tokens, latent video tokens (no measurable gain at LBM scale), CoT-at-inference (degrades).
flowchart TB
P1["Phase 1 β pure co-training (BS=384)<br/>VL : Lang-anno-TRI-Ramen : Lang-anno-human-video<br/>= 1 : 1 : 1<br/>(no continuous robot actions yet)"]
P2["Phase 2 β joint specialization (BS=256)<br/>TRI-Ramen : VL : OXE-Ramen : Lang-anno-human<br/>= 4 : 1 : 4 : 1<br/>(introduce target + cross-embodiment continuous actions)"]
P3["Phase 3 β target specialization (BS=128)<br/>TRI-Ramen : VL : Lang-anno-human<br/>= 9 : 0.5 : 0.5<br/>(drop OXE; keep CE language regularizers)"]
P1 --> P2 --> P3
classDef ph fill:#fff9c4,stroke:#f57f17,color:#000
class P1,P2,P3 ph
Reading the schedule:
- Phase 1 (warm-start the VL side). No continuous-action robot data yet β only language streams. This builds spatial/semantic priors in the VLM half before action gradients touch it.
- Phase 2 (introduce robot continuous actions). Target robot + cross-embodiment robot at 4:4 with VL + human-video captions kept at 1:1 as regularizers. This is the only phase where OXE's morphology gradients enter.
- Phase 3 (drop cross-embodiment, specialize on target). Robot:VL:human-caption = 9 : 0.5 : 0.5 β the 9:1 robot-to-co-training ratio established by the ablation in Fig. S3.
For other "special" policies (latent-action three-phase, FAST-only, etc.), Table S1 ratios are:
| Policy | Phase 1 | Phase 2 | Phase 3 |
|---|---|---|---|
| Three-phase latent-action | TRI-Ramen : OXE-Ramen : Human videos = 3 : 3 : 4, BS=256 | TRI-Ramen : OXE-Ramen = 6 : 4, BS=128 | TRI-Ramen : OXE-Ramen = 9 : 1, BS=128 |
| VL + TRI-OXE-Ramen (FAST) | TRI-Ramen-FAST : VL : OXE-Ramen-FAST = 5 : 2 : 3, BS=128 | TRI-Ramen : VL = 9 : 1, BS=128 | β |
| +Robot-Annotation-Data | VL : TRI-Ramen-LangAnno = 1 : 1, BS=256 | TRI-Ramen : VL = 9 : 1, BS=128 | β |
| +Human-Annotation-Data | VL : TRI-Ramen-LangAnno : Human-LangAnno = 1 : 1 : 1, BS=384 | TRI-Ramen : VL : Human-LangAnno = 9 : 0.5 : 0.5, BS=128 | β |
| +Cross-Embodiment (Final) | VL : TRI-Ramen-LangAnno : Human-LangAnno = 1 : 1 : 1, BS=384 | TRI-Ramen : VL : OXE-Ramen : Human-LangAnno = 4 : 1 : 4 : 1, BS=256 | TRI-Ramen : VL : Human-LangAnno = 9 : 0.5 : 0.5, BS=128 |
| Hyperparameter | Value |
|---|---|
| Optimizer | AdamW (implicit; standard PaliGemma2 recipe) |
| Peak learning rate | 2e-5 |
| LR warm-up steps | 1,000 |
| LR warm-up scheduler | linear |
| LR decay scheduler | cosine decay 2e-5 β 2e-6 over the last 60k steps of each phase |
| Steps per phase (default) | 200,000 |
| Default batch size | 128 (256 or 384 in specific phases per Table S1) |
| Single-task fine-tune steps | 30,000 (cosine decay 2e-5 β 2e-6 over last 24k steps) |
| Hardware | 16 Γ NVIDIA H100 |
| Wall-clock per phase | β 64 hours on 16 H100s @ BS=128 |
| Image observation horizon | 1 (no temporal stacking) |
| Cameras (TRI-Ramen) | 4 (2 scene + 1 wrist per arm) |
| Image augmentation | Random crop 256Γ342 β 224Γ224; color jitter (brightness 0.3, contrast 0.4, saturation 0.5, hue 0.05) |
| Action chunk horizon H | 16 (relative EE pose + gripper width) |
| Inference | Predict 16-step chunk β execute first 8 in open loop β temporal-ensemble across 4 most recent chunks (uniform average); 0.146 s avg latency |
Architecture-comparison note (Fig. S1): Ο0-equivalent and Ο0.5-equivalent variants use BS=112 for 230k steps (higher memory); LBM (this paper) and ΟFAST-equivalent use BS=128 for 200k steps.
The combined objective (Eq. 3 of the paper):
where:
-
$\mathcal{L}_{FM}$ β flow-matching MSE on continuous actions (target robot + cross-embodiment robot) -
$\mathcal{L}_{CE}$ β cross-entropy on language tokens / discrete action tokens / latent action tokens -
$M_{FM}, M_{CE}$ β per-sample masks selecting which loss applies to which batch item -
$w = 0.02$ β the CE loss weight; tuned in Appendix 1B / Fig. S2
Why w = 0.02? Fig. S2 ablation under single-phase VL co-training:
-
$w$ too high β VLM gradient dominates β in-distribution action performance degrades -
$w$ too low β no generalization gain from co-training - 0.02 is the empirical sweet spot
Robot-to-co-training ratio sweet spot (Fig. S3):
- Higher robot ratio β better in-distribution, worse generalization
- Lower robot ratio β better generalization, worse in-distribution
- 9 : 1 (robot : co-training) chosen for the unpaired modalities (VL, OXE, human videos) in the specialization phase
- 6 : 4 (target : OXE) used during phase-1 cross-embodiment mixing β both contribute to the same flow-matching objective so the higher OXE share is tolerable
| Head | Tokens / chunk | Codebook / vocab | Notes |
|---|---|---|---|
| FAST (Pertsch et al.) | 42.1 avg per 16-step chunk | 2,048 | Off-the-shelf β fine-tuning on TRI-Ramen yielded 36.8 tokens but no perf gain (reconstruction error 2.26e-4 either way) |
| VQ-VAE (this paper) | 8 per chunk | 32 | 1D conv encoder + residual blocks; reconstruction + quantization + commitment loss |
| Latent Action Model (LAM) | 8 per video chunk | 32 | 12-layer spatial-temporal IDM transformer + 12-layer spatial FDM + lightweight ActionFDM convolutional decoder |
| GPT-5 captions | language tokens | PaliGemma2 vocab | CE on standard LM head |
LAM training (when used):
- Batch size 1024 on 16 H100s
- 300k steps (~68 hours)
- LR 1e-4 with 1k-step linear warmup and 20k-step cosine decay
- Trained on the full 2,271.6 h human-video corpus
If you only have a fixed budget and want the Final Model:
-
Architecture: PaliGemma2-3B-PT (
google/paligemma2-3b-pt-224, trained jointly β not frozen) + 8-layer ActionFT flow-matching transformer, single observation token aggregated from last 4 VLM layers via adaLN. - Phase 1 (200k steps, BS=384, 16 H100 β 64 h): VL : TRI-Ramen-LangAnno : Human-LangAnno = 1 : 1 : 1. Pure CE β no continuous actions.
- Phase 2 (200k steps, BS=256, β 64 h): TRI-Ramen : VL : OXE-Ramen : Human-LangAnno = 4 : 1 : 4 : 1. Introduce flow matching + keep CE.
- Phase 3 (200k steps, BS=128, β 64 h): TRI-Ramen : VL : Human-LangAnno = 9 : 0.5 : 0.5. Drop OXE.
- Loss: $\mathcal{L} = M_{FM}\mathcal{L}{FM} + 0.02 \cdot M{CE}\mathcal{L}_{CE}$ throughout.
- Optimizer: AdamW; LR 2e-5 with 1,000-step linear warmup; cosine decay 2e-5 β 2e-6 over the last 60k of each 200k phase.
- Total compute: β 192 H100-GPU-days (3 phases Γ 64 h Γ 16 H100s) for the multi-task pretrain; +~10 H100-GPU-days per single-task fine-tune (30k steps).
The paper is admirably explicit about its scope:
- One backbone (PaliGemma2-3B-PT, trained jointly). The paper's verdicts may shift for a much larger backbone or a different VLM family.
- One action head (8-layer ActionFT, flow matching). Does not test, e.g., discrete-diffusion action heads (ICLR-2026-Discrete-Diffusion-VLA) or AR action heads (OpenVLA-class).
- Eval is mostly bimanual Franka tabletop. Mobile manipulation (Ο0.5), humanoid (GR00T), and dexterous-hand (UniHM) are not tested.
- VLM is trained jointly (not frozen) β the co-training-vs-VLM-forgetting result is precisely an artifact of the VLM weights moving: robot-only training drifts them away from the pretrained VL optimum, and VL co-training pulls them back. KI's gradient-routing argument is the alternative design point (insulate the VLM from the continuous-action gradient) that this paper does not adopt.
- Caption-generator quality (GPT-5) is the upper bound on the gain from caption-style modalities. Different vision-language captioners would change the absolute numbers.
- No long-horizon multi-step planning eval at the level where CoT would actually be expected to pay. The CoT-doesn't-help conclusion is therefore a verdict for direct perception-to-action manipulation, not a universal claim. This is the single most likely place where follow-up work will find a counterexample.
What the field still needs:
- The same controlled study at unfrozen-VLM scale.
- The same controlled study with discrete-diffusion action heads β to ask whether the action-head class changes which co-training modalities help.
- A planning-heavy eval suite (CALVIN long-horizon, kitchen-style) to test whether CoT-at-inference pays in environments where multi-step reasoning is genuinely required.
- Per-paper page: LBM Co-training Study
- Knowledge Insulation β the gradient-routing companion
- Hybrid Training β same conclusion on CoT, by curriculum
- Actions as Language β language-relabeling lever
- Human-Video Pretraining β adjacent on human-video data
- VLM4VLA β VLM-benchmark vs VLA-perf disambiguation
- FASTER β discrete-token tokenizer this paper finds unhelpful at LBM scale
- VLA Architectures review
- VLMβAction Connection review β KI vs. cross-attention vs. shared-stack tradeoffs
- Ο0.6 long-form review Β· Ο0.7 long-form review β the production baseline using KI
Verdict: co-training is the best-measured question in the field β VL and cross-embodiment data compound; discrete action tokens don't; mixing must be joint, not sequential. Ratios remain art.
From folklore to measurement: this page's ICLR 2026 study β the RSS 2026 89-policy sequel (4,000 h, 50M VL samples, 58k sim + 2,835 real rollouts) β independent confirmations from the Qwen program and VLM4VLA.
| Claim | Status | Evidence |
|---|---|---|
| VL co-training improves OOD generalization & language following | Confirmed, 3 labs | LBM Γ2; Qwen-RobotManip (+8.2 pp RT-C2R Hard); RSS study |
| Cross-embodiment robot data transfers | Confirmed (conditional on aligned action space β Review-Cross-Embodiment) | RSS study; RobotManip scaling ablation |
| Effective modalities combine cumulatively | Confirmed | RSS study |
| Human video helps as a modality | Confirmed (gripper-scale, diversity-thresholded) | RSS study; H2R Emergence |
| Discrete robot-action tokens as auxiliary signal | No significant benefit β 3Γ replicated | LBM Γ2 + Qwen omission |
| Discrete latent-action tokens as VLM supervision | Effective β refines the scope of the row above | From Pixels to Tokens (ICML 2026 Oral): systematic comparison finds direct VLM supervision with discrete latent action tokens wins |
| Sequential embodied-VQA fine-tuning | Harmful β co-training must be joint | VLM4VLA (all 7 auxiliary tasks negative) |
- Mixture ratios have no theory (Ξ»=0.1 manipulation vs 1.0 navigation in the Qwen suite, undiscussed).
- Scheduling (single- vs multi-phase) shows effects without explanation.
- Per-sample attribution of co-training gains does not exist.
- Budget-constrained adaptation has its own trade-off: ICML 2026's Escaping the Diversity Trap identifies a coverageβdensity tension (anchor-first, then expand to risky boundaries) that pure mixture-ratio thinking misses.