Review LBM Cotraining - Heungwoo/research GitHub Wiki

In-Depth Review β€” A Systematic Study of Data Modalities and Strategies for Co-training LBMs

Paper: A Systematic Study of Data Modalities and Strategies for Co-training Large Behavior Models for Robot Manipulation Lead author: Fanqi Lin Β· Senior author / corresp.: Jose Barreiros ([email protected]) Co-authors include: Kushal Arora, Jean Mercat, Haruki Nishimura, Paarth Shah, Chen Xu, … Affiliation: Toyota Research Institute (TRI) arXiv: 2602.01067 Β· Feb 1, 2026 Code / data: internal TRI; no public release announced as of Feb 2026

This page sits in the VLA architectures review as a training-recipe contribution and pairs with Knowledge Insulation (the canonical "how to protect the VLM during action training" result) as the what data should you co-train on companion.


1. TL;DR

Toyota Research Institute trains 89 LBM (Large Behavior Model) policies β€” a flow-matching VLA built on PaliGemma2-3B β€” under a controlled 5Γ—3 grid of data modality Γ— training strategy, and runs 58,000 simulation rollouts + 2,835 real-robot rollouts to measure each combination.

Three findings overturn folk wisdom that had been propagating through 2025:

  1. Web vision-language data and cross-embodiment robot data dominate β€” they account for almost all of the +36 pp simulation gain and +45 pp real-world language-following gain.
  2. Discrete action tokens (FAST, VQ-VAE) and latent video tokens do not help at LBM scale β€” possibly hurt β€” despite being widely adopted in 2025 VLAs.
  3. Explicit chain-of-thought conditioning at inference is harmful for direct perception-to-action tasks; the gain people attributed to CoT is from training-time supervision, not inference-time conditioning. This empirically corroborates from a different angle the Hybrid Training thesis.

Two further findings refine the recipe:

  1. Cross-embodiment data is best used in phase 1 only β€” keeping it in phase 2 hurts.
  2. Robot-only training erodes the VLM's MMBench/GQA scores; balanced co-training restores them β€” i.e., VL co-training is partly a regularization-against-forgetting story, not just a generalization-gain story.

2. Why this paper matters in the 2026 landscape

By Q1 2026 the community had three loud co-training claims fighting for default-recipe status:

  • PI / Knowledge Insulation (NeurIPS 2025): the gradient from the action expert must not back-prop into the VLM, or you lose the VLM's pretrained knowledge.
  • OpenVLA / FAST / OmniSAT: discrete action tokenization is the way to share parameters between vision-language and action.
  • ECoT family: chain-of-thought traces (text or latent) attached to actions improve generalization.

The TRI paper is the first controlled comparison that grades all three with the same backbone, the same robot platform, and the same evaluation suite. The verdict for each:

2025 claim TRI 2026 verdict
Co-training with VL data helps βœ… confirmed β€” largest single-modality gain
Discrete action tokenization helps ❌ no significant benefit at LBM scale
Cross-embodiment data helps βœ… β€” but only in phase 1
Chain-of-thought conditioning at inference helps ❌ does not improve and often degrades for direct manipulation
Latent video tokens help ❌ marginal/null at scale; useful only in low-data regime
Robot-only training is fine ❌ erodes the VLM's visiolinguistic ability

This is the empirical companion paper that the field had been waiting for since the Sergey Levine et al. What Matters in Learning from Large-Scale Datasets (ICRA 2025) line β€” same spirit, but at LBM scale and including the language/CoT axes that the older paper did not cover.


3. Architecture under study

flowchart LR
  IMG[Multi-camera observation] --> VLM
  LANG[Language instruction] --> VLM
  STATE[Proprio state] --> VLM
  VLM[PaliGemma2-3B-PT VLM<br/>trained jointly with action expert]
  VLM -- single observation token,<br/>extracted from last 4 layers --> COND[adaLN conditioning]
  COND --> AFT[ActionFT β€” 8-layer flow-matching transformer]
  AFT --> ACT[Action chunk: 16 steps<br/>relative EE pose + gripper width]

  VLM -.discrete-token CE losses.-> AUX[Auxiliary discrete heads<br/>FAST tokens / VQ-VAE / language captions]

  classDef vlm fill:#bbdefb,stroke:#1565c0,color:#000
  classDef act fill:#c8e6c9,stroke:#2e7d32,color:#000
  classDef aux fill:#fff9c4,stroke:#f57f17,color:#000
  class VLM,COND vlm
  class AFT,ACT act
  class AUX aux
Loading
  • Backbone: PaliGemma2-3B-PT, trained jointly with the action expert (not frozen) β€” the VLM "can optionally be trained to generate text or discrete action tokens." This is essential to the paper's catastrophic-forgetting result: only because the VLM weights move can robot-only training erode its MMBench/GQA scores and VL co-training restore them.
  • Action expert: ActionFT, an 8-layer flow-matching transformer, conditioned via adaLN on a single observation token built from the last 4 layers of the VLM.
  • Action representation: continuous flow-matching for the primary head; auxiliary discrete heads added under specific co-training settings.
  • Action horizon: 16-step chunks.

Loss = flow-matching MSE on continuous actions + cross-entropy on auxiliary token / language streams when present, weighted by computed masks.

Note (correction): an earlier version of this page described the VLM as "frozen during action training." The source describes a jointly-trained end-to-end system; the VLM is not frozen.

This architecture is deliberately close to the public PaliGemma-3B + flow-matching recipe so the conclusions transfer to other LBM-class systems built on the same template (Ο€0.5 / Ο€0.6 / Steerable-Ο€0.5, GR00T-N1.x family modulo backbone size).


4. The five co-training data modalities

flowchart TB
  subgraph M1[Modality 1 β€” Standard VL data Β· 50M samples]
    direction LR
    RP[RoboPoint<br/>1.3M images Β· 8.2M QA]
    RS[RefSpatial<br/>2.5M images Β· 20M QA]
  end
  subgraph M2[Modality 2 β€” Dense robot annotations Β· 523h]
    direction LR
    SC[Scripted heuristics<br/>EE-motion descriptions]
    GC[GPT-5 captions<br/>rich semantic descriptions]
  end
  subgraph M3[Modality 3 β€” Cross-embodiment robot Β· 1,150h]
    OXE[OXE-Ramen<br/>12 robots Β· 924 tasks Β· 466k demos]
  end
  subgraph M4[Modality 4 β€” Human videos Β· 2,271h]
    direction LR
    LAM[Latent Action Model<br/>32 codebook Β· 8 tokens/chunk]
    HC[GPT-5 captions of motion]
  end
  subgraph M5[Modality 5 β€” Discrete robot tokens Β· 523h]
    direction LR
    FAST[FAST tokens<br/>~42 / chunk Β· vocab 2,048]
    VQ[VQ-VAE tokens<br/>8 / chunk Β· codebook 32]
  end

  classDef good fill:#c8e6c9,stroke:#2e7d32,color:#000
  classDef bad fill:#ffcdd2,stroke:#c62828,color:#000
  class M1,M3 good
  class M2 good
  class M5 bad
Loading

Color coding: green = the paper's "effective" verdict, red = "no significant benefit", mixed for M2/M4 (caption variants effective, token variants not).

4.1 Modality 1 β€” Standard vision-language data

  • Sources: RoboPoint (1.3M images, 8.2M QA pairs β€” robotics-relevant pointing and spatial QA) and RefSpatial (2.5M images, 20M QA β€” refer-and-locate spatial reasoning). Combined β‰ˆ 50M VL samples.
  • Loss: standard CE on the VLM language head.
  • Verdict: βœ… largest single-modality contribution to both unseen-task generalization (sim) and language following (real). Cheap to scale because no robot data collection is required.

4.2 Modality 2 β€” Dense language annotations of robot trajectories

  • Source: TRI-Ramen (~523 h of TRI's in-house manipulation demos β€” 403 tasks, 53,411 demonstrations on the dual-Franka platform).
  • Two caption variants:
    • Scripted heuristics β€” derive end-effector motion descriptions from action stream ("gripper moves left and down 5 cm, closes").
    • GPT-5 captions β€” vision-language model emits rich semantic captions for short trajectory segments at ~1–2 second intervals.
  • Verdict: βœ… GPT-5 captions clearly beat scripted heuristics. The semantic richness β€” object names, contact events, intent β€” is what matters. Scripted heuristics give the action a redundant motion description that the VLM does not need.

4.3 Modality 3 β€” Cross-embodiment robot data

  • Source: OXE-Ramen β€” a TRI-curated subset of Open X-Embodiment covering 12 robot setups, 924 tasks, 466k demos, totaling 1,150 h.
  • Used as: native continuous actions (not retokenized).
  • Verdict: βœ… effective, but best used in phase 1 only. Keeping it in phase 2 hurts β€” the morphology gap between the OXE robots and the target dual-Franka platform creates conflicting gradients during specialization.

4.4 Modality 4 β€” Human videos

  • Sources (2,271 h after filtering): Ego4D, EgoDex, Something-Something V2, Epic Kitchen, HoloAssist.
  • Two representations tested:
    • Latent action tokens β€” Latent Action Model (LAM) trained to compress short clips into 8 tokens per chunk over a codebook of 32; supervised via CE.
    • GPT-5 motion captions β€” VLM-generated language descriptions of the human motion.
  • Verdict: βœ… caption variant works. ❌ latent-action variant only helps in the low-data regime; gains evaporate at scale. The paper's framing: human videos contribute visual diversity through language, not through action-token transfer.

4.5 Modality 5 β€” Discrete robot action tokens

  • Source: TRI-Ramen retokenized two ways:
    • FAST β€” DCT-based, β‰ˆ 42 tokens per 16-step chunk over a 2,048 vocab.
    • VQ-VAE β€” 8 tokens per chunk over a 32 codebook.
  • Used as: auxiliary CE prediction targets alongside the primary flow-matching head.
  • Verdict: ❌ no significant benefit; FAST sometimes hurts unseen-task performance. This is one of the paper's most important findings β€” Knowledge Insulation NeurIPS-2025-Knowledge-Insulation uses FAST tokens to give the VLM a CE signal, but TRI's controlled study shows the signal itself is not what matters; KI's gain must come from the gradient routing, not from the token target.

5. The three training strategies

flowchart LR
  subgraph SP[Single-phase joint]
    SP1[Mix target + co-training<br/>train one phase end-to-end]
  end
  subgraph TP1[Two-phase Β· 1st-phase only]
    TP1a[Phase 1: co-training data only]
    TP1b[Phase 2: target robot continuous actions only]
    TP1a --> TP1b
  end
  subgraph TP2[Two-phase Β· full co-training]
    TP2a[Phase 1: co-training data only]
    TP2b[Phase 2: target + co-training jointly]
    TP2a --> TP2b
  end

  classDef ph fill:#fff9c4,stroke:#f57f17,color:#000
  class SP1,TP1a,TP1b,TP2a,TP2b ph
Loading

The strategy that wins depends on the modality:

Modality Best strategy
Standard VL Two-phase full (keep VL in phase 2)
Robot annotations (GPT-5) Two-phase, 1st-phase only
Cross-embodiment robot Two-phase, 1st-phase only
Human-video captions Two-phase full (keep in phase 2)
FAST / VQ-VAE / latent video tokens none of the strategies recover usable gain

Pattern: language-bearing modalities want to be co-trained throughout (phase 1 + phase 2). Action-stream modalities from non-target morphologies want to be dropped after phase 1 so the policy can specialize.

This pattern is consistent with how the Ο€-series uses Knowledge Insulation: the VLM keeps getting language-side gradients (CE on tokens / captions) throughout training, while the action expert specializes β€” TRI's empirical result is the data-side counterpart of KI's gradient-routing result.


6. Evaluation suite

6.1 Simulation β€” Drake-based

  • Tasks: 13 seen + 8 unseen (semantic / multi-step / compositional generalization).
  • Conditions: nominal + four distribution shifts β€” lighting, textures, camera viewpoint, backgrounds.
  • Rollouts: 50 per task per condition β†’ 58,000+ total.

6.2 Real-world β€” dual Franka Panda

flowchart LR
  subgraph LF[Language-following Β· 45 rollouts/condition]
    LF1[Seen objects from training]
    LF2[Instruction generalization<br/>paraphrases Β· category Β· attribute]
    LF3[Unseen objects Β· 52 novel]
  end
  subgraph LH[Long-horizon dexterous Β· 30 rollouts each Β· ~93 s Β· 13 steps avg]
    LH1[PackItemsIntoStringBag<br/>fine-grained bottle capping]
    LH2[PourIngredientsIntoSoup<br/>scooping with spatula]
    LH3[StoreCleanDishes<br/>transparent objects]
  end
Loading
  • Total real-robot rollouts: 2,835.

6.3 VLM-benchmark probe

In addition to manipulation, the LBM's PaliGemma2 backbone is evaluated on MMBench, GQA, and other standard VL benchmarks before/after each training regime β€” to detect catastrophic forgetting of visiolinguistic capability.


7. Headline numbers

7.1 Simulation, unseen-task average

Setting Unseen tasks
Baseline (target-only training) β‰ˆ 36.2%
+ Standard VL ↑↑
+ GPT-5 robot captions ↑
+ Cross-embodiment (phase 1 only) ↑
+ Human-video captions ↑
All effective modalities stacked 72.6%

β†’ +36.4 percentage points from the right co-training mix; gains stack additively rather than saturating.

7.2 Real-world language following

Setting Task completion (avg over 3 conditions)
Baseline β‰ˆ 24.1%
All effective modalities 69.4%

β†’ +45.3 pp. The unseen-objects axis is the largest single beneficiary β€” VL data + human-video captions push the model to recognize objects it never saw on the target robot.

7.3 Fine-tune-to-new-task

After pretraining, fine-tune on 200 demonstrations of three unseen long-horizon tasks (bottle-capping in bag, scooping into soup, storing transparent dishes):

Setting Completion
Single-task untrained baseline 47.3%
Pretrained baseline (no co-training) β†’ fine-tune 67.4%
Full co-trained LBM β†’ fine-tune 90.2%

The +22.8 pp delta over the same fine-tune procedure on the no-co-training baseline is the cleanest demonstration in the paper that co-training learns transferable representations, not just multi-task amortization.

7.4 VLM-benchmark preservation

Robot-only training drops PaliGemma2 MMBench / GQA scores measurably vs. the original VLM weights. Adding standard VL co-training restores them. This frames part of the gain as preventing catastrophic forgetting rather than learning new capability.


8. The chain-of-thought conditioning experiment

This is the section that will be most cited because it directly contradicts the prevailing 2024-2025 ECoT narrative.

8.1 Setup

Three CoT trace types are tested:

  • Scripted captions (Modality 2a)
  • VLM-generated captions (Modality 2b / Modality 4b)
  • Latent action tokens (Modality 4a)

For each, three CoT strategies:

# Train Inference
1 50% with CoT, 50% without Without CoT (skip generation)
2 50% with CoT, 50% without With CoT β€” generate trace then condition
3 Always with CoT Always with CoT

8.2 Result

All three CoT inference strategies either match or degrade the implicit two-phase co-training baseline.

The paper's diagnosis:

  • The eval tasks are direct perception-to-action (pick-and-place, simple long-horizon scoops). They do not require multi-step planning.
  • VLM-generated captions hallucinate at non-trivial rates; latent action tokens contain reconstruction noise.
  • When the action prediction is conditioned on the noisy generated CoT, errors propagate: a hallucinated subgoal directly distorts the action chunk that follows.
  • When CoT is only used as an auxiliary CE training target and skipped at inference (option 1), the model still doesn't gain β€” the implicit co-training already captures whatever benefit was available.

8.3 Implication for the field

This is the paper's most direct dialogue with the broader 2026 reasoning-augmented VLA literature:

  • Hybrid Training had argued: the ECoT generalization gain is from training-time supervision; drop the reasoning at inference. TRI's controlled experiment empirically validates that thesis for the LBM-scale setting.
  • Actions as Language had argued: subtask relabeling is the lever. TRI confirms: language captions of robot/human trajectories are the most valuable form of co-training, but only as CE training targets, not as inference conditioning.
  • Embodied-R1 and dVLA are the two papers that do make CoT-at-inference work β€” they pay either via RL (E-R1) or via parallel decoding (dVLA). TRI's paper sharpens their justification: vanilla CoT-at-inference doesn't pay off; you need a specific architectural or training reason to use it.

9. Why "no benefit from discrete action tokens" matters

This is the second most important finding for the field's recipes.

The 2024–2025 consensus had been that discrete action tokens (FAST, BPE, VQ-VAE) are the way to:

  1. Share parameters with VLM language heads (one CE loss).
  2. Make actions look like language so that a frozen LM can be co-trained.
  3. Compress high-frequency action streams.

Knowledge Insulation β€” the foundation of Ο€0.6/Ο€0.7 β€” uses FAST tokens specifically as the CE training target for the VLM half, while the continuous flow-matching head remains the action head.

TRI's controlled experiment shows that, with this exact architecture (PaliGemma2 + flow-matching + auxiliary FAST CE head), the FAST CE loss adds no measurable benefit and sometimes hurts unseen-task performance.

Two ways to reconcile:

  1. The KI gain comes from gradient routing, not from the token target. That is, what matters is that the VLM's gradient is insulated from the continuous head β€” the auxiliary CE target was a vehicle for that, not the cause of the gain.
  2. The benefit of FAST tokens depends on the eval suite. Possible β€” TRI's eval is dual-Franka manipulation; PI's flagship evals include very different domains (mobile manipulation, laundry folding). The TRI paper does not claim a universal verdict, only the LBM-scale verdict on its specific eval suite.

Either way, this is now an open empirical question the field needs to settle in 2026 β€” not a closed-recipe assumption.


10. Recipe for practitioners

Distilled from Β§III of the paper:

  1. Always include large-scale standard VL data (RoboPoint + RefSpatial or equivalent). It gives the largest single-modality gain and prevents VLM forgetting. Cheap; no robot data collection.
  2. Always include GPT-5-generated captions over your robot demos. Better than scripted heuristics. Keep them as a CE target throughout phase 2.
  3. Include cross-embodiment robot data (OXE) only in phase 1. Drop it for specialization in phase 2.
  4. Include human-video captions (not latent action tokens) and keep them throughout both phases.
  5. Skip discrete action tokenization (FAST / VQ-VAE) as an auxiliary head unless you have a specific reason (low-data regime, KI-style gradient routing argument).
  6. Skip explicit chain-of-thought conditioning at inference. If you want CoT's training-time benefit, use it as a CE auxiliary loss only and predict actions without generating CoT at inference.
  7. Two-phase training is the default. Single-phase joint training works too but does not match the two-phase ceiling.

11. Full training recipe β€” hyperparameters, ratios & loss weights

This section consolidates everything needed to reproduce the Final Model ("+Cross-Embodiment-Robot-Data" row in Table S1) β€” the recipe that achieved the best generalization in the paper. Numbers are pulled from Appendix 2 (Training Details), Table S1 (per-phase data ratios), Table S2 (shared hyperparameters), and Table S3 (human-video data composition).

Note on "two-phase" vs "three-phase": The paper's main text (Fig. 9) reports that the best single-modality policies are either two-phase full co-training (standard VL, human-video annotations) or two-phase 1st-phase-only (scripted/VLM robot annotations, cross-embodiment data). That statement is about isolated modalities. The cumulative Final Model (Β§III-C, "+Cross-Embodiment-Robot-Data") is a separate model that stacks all effective modalities, and in Table S1 it has populated 1st-/2nd-/3rd-phase columns β€” i.e. it is genuinely three-phase. The two are not in conflict: "two-phase" is the per-modality finding; "three-phase" is the stacked-recipe schedule. (Audit 2026-06: all numbers in this section were cross-checked against the arXiv HTML full text incl. appendix and confirmed β€” see report.)

11.1 Optimal data composition for the Final Model

The best-performing policy is a three-phase schedule that mixes five modalities at carefully tuned ratios:

Modality Volume Role
TRI-Ramen (target dual-Franka teleop) 523 h Continuous flow-matching anchor β€” the "golden" target distribution
OXE-Ramen (cross-embodiment) 1,150 h, 466k demos, 12 robots, 924 tasks Continuous flow-matching, phase-1+2 only (dropped in phase 3)
Standard VL (RoboPoint 8.2M QA + RefSpatial 20M QA) β‰ˆ 50M VL samples CE on language head; throughout all phases
GPT-5 language annotations of TRI-Ramen β‰ˆ 523 h Γ— 1–2 captions/s CE on language head; throughout all phases
GPT-5 language annotations of human videos (Ego4D 774.5 h, EgoDex 744.4 h + reversed 455.7 h, Sth-Sth-V2 155.8 h, Epic Kitchen 60.4 h, HoloAssist 80.8 h β‰ˆ 2,271.6 h) 9.0M annotations across 2,271.6 h CE on language head; throughout all phases

Excluded by ablation: FAST tokens, VQ-VAE tokens, latent video tokens (no measurable gain at LBM scale), CoT-at-inference (degrades).

11.2 The three-phase training schedule (Final Model)

flowchart TB
  P1["Phase 1 β€” pure co-training (BS=384)<br/>VL : Lang-anno-TRI-Ramen : Lang-anno-human-video<br/>= 1 : 1 : 1<br/>(no continuous robot actions yet)"]
  P2["Phase 2 β€” joint specialization (BS=256)<br/>TRI-Ramen : VL : OXE-Ramen : Lang-anno-human<br/>= 4 : 1 : 4 : 1<br/>(introduce target + cross-embodiment continuous actions)"]
  P3["Phase 3 β€” target specialization (BS=128)<br/>TRI-Ramen : VL : Lang-anno-human<br/>= 9 : 0.5 : 0.5<br/>(drop OXE; keep CE language regularizers)"]
  P1 --> P2 --> P3
  classDef ph fill:#fff9c4,stroke:#f57f17,color:#000
  class P1,P2,P3 ph
Loading

Reading the schedule:

  • Phase 1 (warm-start the VL side). No continuous-action robot data yet β€” only language streams. This builds spatial/semantic priors in the VLM half before action gradients touch it.
  • Phase 2 (introduce robot continuous actions). Target robot + cross-embodiment robot at 4:4 with VL + human-video captions kept at 1:1 as regularizers. This is the only phase where OXE's morphology gradients enter.
  • Phase 3 (drop cross-embodiment, specialize on target). Robot:VL:human-caption = 9 : 0.5 : 0.5 β€” the 9:1 robot-to-co-training ratio established by the ablation in Fig. S3.

For other "special" policies (latent-action three-phase, FAST-only, etc.), Table S1 ratios are:

Policy Phase 1 Phase 2 Phase 3
Three-phase latent-action TRI-Ramen : OXE-Ramen : Human videos = 3 : 3 : 4, BS=256 TRI-Ramen : OXE-Ramen = 6 : 4, BS=128 TRI-Ramen : OXE-Ramen = 9 : 1, BS=128
VL + TRI-OXE-Ramen (FAST) TRI-Ramen-FAST : VL : OXE-Ramen-FAST = 5 : 2 : 3, BS=128 TRI-Ramen : VL = 9 : 1, BS=128 β€”
+Robot-Annotation-Data VL : TRI-Ramen-LangAnno = 1 : 1, BS=256 TRI-Ramen : VL = 9 : 1, BS=128 β€”
+Human-Annotation-Data VL : TRI-Ramen-LangAnno : Human-LangAnno = 1 : 1 : 1, BS=384 TRI-Ramen : VL : Human-LangAnno = 9 : 0.5 : 0.5, BS=128 β€”
+Cross-Embodiment (Final) VL : TRI-Ramen-LangAnno : Human-LangAnno = 1 : 1 : 1, BS=384 TRI-Ramen : VL : OXE-Ramen : Human-LangAnno = 4 : 1 : 4 : 1, BS=256 TRI-Ramen : VL : Human-LangAnno = 9 : 0.5 : 0.5, BS=128

11.3 Shared training hyperparameters (Table S2 + Appendix 2)

Hyperparameter Value
Optimizer AdamW (implicit; standard PaliGemma2 recipe)
Peak learning rate 2e-5
LR warm-up steps 1,000
LR warm-up scheduler linear
LR decay scheduler cosine decay 2e-5 β†’ 2e-6 over the last 60k steps of each phase
Steps per phase (default) 200,000
Default batch size 128 (256 or 384 in specific phases per Table S1)
Single-task fine-tune steps 30,000 (cosine decay 2e-5 β†’ 2e-6 over last 24k steps)
Hardware 16 Γ— NVIDIA H100
Wall-clock per phase β‰ˆ 64 hours on 16 H100s @ BS=128
Image observation horizon 1 (no temporal stacking)
Cameras (TRI-Ramen) 4 (2 scene + 1 wrist per arm)
Image augmentation Random crop 256Γ—342 β†’ 224Γ—224; color jitter (brightness 0.3, contrast 0.4, saturation 0.5, hue 0.05)
Action chunk horizon H 16 (relative EE pose + gripper width)
Inference Predict 16-step chunk β†’ execute first 8 in open loop β†’ temporal-ensemble across 4 most recent chunks (uniform average); 0.146 s avg latency

Architecture-comparison note (Fig. S1): Ο€0-equivalent and Ο€0.5-equivalent variants use BS=112 for 230k steps (higher memory); LBM (this paper) and Ο€FAST-equivalent use BS=128 for 200k steps.

11.4 Loss formulation & weights

The combined objective (Eq. 3 of the paper):

$$\mathcal{L} ;=; M_{FM} \cdot \mathcal{L}_{FM} ;+; w \cdot M_{CE} \cdot \mathcal{L}_{CE}$$

where:

  • $\mathcal{L}_{FM}$ β€” flow-matching MSE on continuous actions (target robot + cross-embodiment robot)
  • $\mathcal{L}_{CE}$ β€” cross-entropy on language tokens / discrete action tokens / latent action tokens
  • $M_{FM}, M_{CE}$ β€” per-sample masks selecting which loss applies to which batch item
  • $w = 0.02$ β€” the CE loss weight; tuned in Appendix 1B / Fig. S2

Why w = 0.02? Fig. S2 ablation under single-phase VL co-training:

  • $w$ too high β†’ VLM gradient dominates β†’ in-distribution action performance degrades
  • $w$ too low β†’ no generalization gain from co-training
  • 0.02 is the empirical sweet spot

Robot-to-co-training ratio sweet spot (Fig. S3):

  • Higher robot ratio β†’ better in-distribution, worse generalization
  • Lower robot ratio β†’ better generalization, worse in-distribution
  • 9 : 1 (robot : co-training) chosen for the unpaired modalities (VL, OXE, human videos) in the specialization phase
  • 6 : 4 (target : OXE) used during phase-1 cross-embodiment mixing β€” both contribute to the same flow-matching objective so the higher OXE share is tolerable

11.5 Auxiliary head specifications

Head Tokens / chunk Codebook / vocab Notes
FAST (Pertsch et al.) 42.1 avg per 16-step chunk 2,048 Off-the-shelf β€” fine-tuning on TRI-Ramen yielded 36.8 tokens but no perf gain (reconstruction error 2.26e-4 either way)
VQ-VAE (this paper) 8 per chunk 32 1D conv encoder + residual blocks; reconstruction + quantization + commitment loss
Latent Action Model (LAM) 8 per video chunk 32 12-layer spatial-temporal IDM transformer + 12-layer spatial FDM + lightweight ActionFDM convolutional decoder
GPT-5 captions language tokens PaliGemma2 vocab CE on standard LM head

LAM training (when used):

  • Batch size 1024 on 16 H100s
  • 300k steps (~68 hours)
  • LR 1e-4 with 1k-step linear warmup and 20k-step cosine decay
  • Trained on the full 2,271.6 h human-video corpus

11.6 Distilled "do-this" recipe for reproduction

If you only have a fixed budget and want the Final Model:

  1. Architecture: PaliGemma2-3B-PT (google/paligemma2-3b-pt-224, trained jointly β€” not frozen) + 8-layer ActionFT flow-matching transformer, single observation token aggregated from last 4 VLM layers via adaLN.
  2. Phase 1 (200k steps, BS=384, 16 H100 β‰ˆ 64 h): VL : TRI-Ramen-LangAnno : Human-LangAnno = 1 : 1 : 1. Pure CE β€” no continuous actions.
  3. Phase 2 (200k steps, BS=256, β‰ˆ 64 h): TRI-Ramen : VL : OXE-Ramen : Human-LangAnno = 4 : 1 : 4 : 1. Introduce flow matching + keep CE.
  4. Phase 3 (200k steps, BS=128, β‰ˆ 64 h): TRI-Ramen : VL : Human-LangAnno = 9 : 0.5 : 0.5. Drop OXE.
  5. Loss: $\mathcal{L} = M_{FM}\mathcal{L}{FM} + 0.02 \cdot M{CE}\mathcal{L}_{CE}$ throughout.
  6. Optimizer: AdamW; LR 2e-5 with 1,000-step linear warmup; cosine decay 2e-5 β†’ 2e-6 over the last 60k of each 200k phase.
  7. Total compute: β‰ˆ 192 H100-GPU-days (3 phases Γ— 64 h Γ— 16 H100s) for the multi-task pretrain; +~10 H100-GPU-days per single-task fine-tune (30k steps).

12. Limitations & open questions

The paper is admirably explicit about its scope:

  • One backbone (PaliGemma2-3B-PT, trained jointly). The paper's verdicts may shift for a much larger backbone or a different VLM family.
  • One action head (8-layer ActionFT, flow matching). Does not test, e.g., discrete-diffusion action heads (ICLR-2026-Discrete-Diffusion-VLA) or AR action heads (OpenVLA-class).
  • Eval is mostly bimanual Franka tabletop. Mobile manipulation (Ο€0.5), humanoid (GR00T), and dexterous-hand (UniHM) are not tested.
  • VLM is trained jointly (not frozen) β€” the co-training-vs-VLM-forgetting result is precisely an artifact of the VLM weights moving: robot-only training drifts them away from the pretrained VL optimum, and VL co-training pulls them back. KI's gradient-routing argument is the alternative design point (insulate the VLM from the continuous-action gradient) that this paper does not adopt.
  • Caption-generator quality (GPT-5) is the upper bound on the gain from caption-style modalities. Different vision-language captioners would change the absolute numbers.
  • No long-horizon multi-step planning eval at the level where CoT would actually be expected to pay. The CoT-doesn't-help conclusion is therefore a verdict for direct perception-to-action manipulation, not a universal claim. This is the single most likely place where follow-up work will find a counterexample.

What the field still needs:

  • The same controlled study at unfrozen-VLM scale.
  • The same controlled study with discrete-diffusion action heads β€” to ask whether the action-head class changes which co-training modalities help.
  • A planning-heavy eval suite (CALVIN long-horizon, kitchen-style) to test whether CoT-at-inference pays in environments where multi-step reasoning is genuinely required.

13. Pointers to related pages


πŸ—“ State of the Field (updated Aug 2026)

Verdict: co-training is the best-measured question in the field β€” VL and cross-embodiment data compound; discrete action tokens don't; mixing must be joint, not sequential. Ratios remain art.

πŸ“ˆ Trend

From folklore to measurement: this page's ICLR 2026 study β†’ the RSS 2026 89-policy sequel (4,000 h, 50M VL samples, 58k sim + 2,835 real rollouts) β†’ independent confirmations from the Qwen program and VLM4VLA.

βœ… Consolidated verdicts

Claim Status Evidence
VL co-training improves OOD generalization & language following Confirmed, 3 labs LBM Γ—2; Qwen-RobotManip (+8.2 pp RT-C2R Hard); RSS study
Cross-embodiment robot data transfers Confirmed (conditional on aligned action space β€” Review-Cross-Embodiment) RSS study; RobotManip scaling ablation
Effective modalities combine cumulatively Confirmed RSS study
Human video helps as a modality Confirmed (gripper-scale, diversity-thresholded) RSS study; H2R Emergence
Discrete robot-action tokens as auxiliary signal No significant benefit β€” 3Γ— replicated LBM Γ—2 + Qwen omission
Discrete latent-action tokens as VLM supervision Effective β€” refines the scope of the row above From Pixels to Tokens (ICML 2026 Oral): systematic comparison finds direct VLM supervision with discrete latent action tokens wins
Sequential embodied-VQA fine-tuning Harmful β€” co-training must be joint VLM4VLA (all 7 auxiliary tasks negative)

⚠️ Limitations & open problems

  • Mixture ratios have no theory (Ξ»=0.1 manipulation vs 1.0 navigation in the Qwen suite, undiscussed).
  • Scheduling (single- vs multi-phase) shows effects without explanation.
  • Per-sample attribution of co-training gains does not exist.
  • Budget-constrained adaptation has its own trade-off: ICML 2026's Escaping the Diversity Trap identifies a coverage–density tension (anchor-first, then expand to risky boundaries) that pure mixture-ratio thinking misses.

← Back to Home Β· ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️