ICLR 2026 LBM Cotraining - Heungwoo/research GitHub Wiki

LBM Co-training Study β€” Which Data Modalities Actually Help a Large Behavior Model?

Venue: arXiv preprint (Feb 1, 2026) β€” Toyota Research Institute (TRI) arXiv: 2602.01067 Category: Training Approach / Data Strategy (empirical study) Trend tag: Trend 2 (training recipes) Β· Trend 4 (data scaling) Scale of study: 4,000 hours of robot/human manipulation + 50M VL samples Β· 89 policies trained Β· 58,000 sim rollouts + 2,835 real-robot rollouts

Approach diagram

flowchart LR
  subgraph Mod[5 co-training modalities]
    M1[1. Standard VL data<br/>RoboPoint + RefSpatial<br/>50M samples]
    M2[2. Dense robot annotations<br/>scripted + GPT-5 captions<br/>523h]
    M3[3. Cross-embodiment robot<br/>OXE-Ramen 1,150h<br/>12 robots / 924 tasks]
    M4[4. Human videos<br/>2,271h Ego4D / EgoDex / SSv2<br/>latent actions OR captions]
    M5[5. Discrete robot tokens<br/>FAST or VQ-VAE<br/>523h]
  end

  subgraph LBM[VLA / LBM under test]
    VLM[PaliGemma2-3B-PT<br/>vision-language backbone]
    AH[ActionFT<br/>8-layer flow-matching transformer]
    VLM -- adaLN conditioning --> AH
  end

  subgraph Strat[3 training strategies]
    S1[Single-phase joint]
    S2[2-phase: 1st-phase only]
    S3[2-phase: full co-training]
  end

  Mod --> LBM
  Strat --> LBM
  LBM --> Eval[Sim 21 tasks Β· Real 9+ tasks<br/>distribution shift / unseen / language following]
Loading

Problem

Large Behavior Models (LBMs) β€” multi-task imitation policies that scale to dexterous manipulation β€” are bottlenecked by insufficient robot data coverage. The community's response is co-training with heterogeneous data (web VL, OXE cross-embodiment, human videos, action tokens, …), but until now the field has had no controlled comparison of which modalities and which training schedules actually pay off.

Concretely the paper attacks four open questions:

  1. Which co-training data modalities help an LBM, and how much?
  2. Should the modality be used in the first phase only (pretraining), throughout (joint co-training), or in a single-phase mix?
  3. Does chain-of-thought conditioning on co-training-derived traces improve action prediction?
  4. Are gains additive when modalities are stacked?

Method

Backbone (LBM under study). PaliGemma2-3B-PT VLM (google/paligemma2-3b-pt-224, fine-tuned during co-training β€” its trainability is exactly what makes the MMBench/GQA preservation finding possible) feeds an 8-layer ActionFT flow-matching transformer. A single global conditioning token is extracted from the VLM's last four layers and injected into every ActionFT layer via an adaLN MLP (alongside the flow-matching timestep). Action chunk = 16 steps (~1.6 s at 10 Hz) of relative end-effector poses w.r.t. the station base frame + gripper widths. A small auxiliary discrete-token head is added when training with discrete-action losses (CE) alongside flow matching (FM).

Five co-training modalities studied:

# Modality Source Format
1 Standard VL RoboPoint (1.3M / 8.2M QA) + RefSpatial (2.5M / 20M QA) VQA, object localization, spatial reasoning
2 Dense robot annotations TRI-Ramen (β‰ˆ523h) with two caption types: (a) scripted-heuristic motion descriptions, (b) GPT-5 generated rich descriptions Per-frame language captions at ~1–2 s intervals
3 Cross-embodiment robot OXE-Ramen (curated Open X-Embodiment) β€” 12 setups, 924 tasks, 466k demos, 1,150h Native continuous actions, diverse morphologies
4 Human videos Ego4D, EgoDex, Something-Something V2, Epic Kitchen, HoloAssist β€” 2,271h filtered. Two representations: (a) latent action tokens via Latent Action Model (codebook 32, 8 tokens/chunk), (b) GPT-5 motion captions Either token sequence (CE loss) or caption (CE)
5 Discrete robot tokens TRI-Ramen retokenized two ways: (a) FAST (~42.1 tokens / chunk, vocab 2,048), (b) VQ-VAE (8 tokens, codebook 32) Compressed action tokens (CE)

Three training strategies:

  1. Single-phase joint β€” target robot data + co-training modalities mixed in one phase.
  2. Two-phase, 1st-phase-only β€” co-training data in phase 1, then target robot continuous actions in phase 2.
  3. Two-phase, full co-training β€” co-training data in phase 1, then target + co-training jointly in phase 2.

Loss. Flow-matching for continuous actions, cross-entropy for any discrete token / language stream, weighted sum with computed masks.

Chain-of-thought conditioning experiment. Three CoT variants tested on three trace types (scripted captions, VLM captions, latent actions): (i) train with CoT 50% of the time, infer without CoT; (ii) same training, generate CoT then condition actions on it at inference; (iii) always condition on CoT during training and inference.

Results

Simulation β€” 21 tasks (13 seen + 8 unseen), 50 rollouts each, 58k+ total rollouts.

Setting Unseen-task success
No co-training (target only) β‰ˆ 36.2%
All effective modalities stacked 72.6% (+36.4 pp)

Distribution-shift robustness (lighting, textures, camera, backgrounds) climbs in lockstep with each effective modality.

Real-world (dual Franka Panda, 2,835 rollouts).

Eval Baseline Final stacked LBM
Language following (45 / cond.) β‰ˆ 24.1% 69.4% (+45.3 pp)
Long-horizon dex tasks fine-tune (200 demos, 30 rollouts) 67.4% (FT'd baseline) / 47.3% (single-task untrained) 90.2%

Per-modality ranking (effective β†’ ineffective):

  1. βœ… Standard VL data β€” largest single-modality gain
  2. βœ… VLM-generated robot annotations β€” > scripted heuristics
  3. βœ… Cross-embodiment robot data β€” most useful in phase-1 only
  4. βœ… VLM-generated human-video annotations β€” sustained co-training helps
  5. ❌ FAST discrete tokens β€” no significant benefit; degrades unseen tasks
  6. ❌ Latent action tokens from human video β€” only in low-data regime, diminishing
  7. ❌ VQ-VAE robot tokens β€” marginal/null

Chain-of-thought conditioning. All three CoT variants either match or degrade vs. implicit two-phase co-training on the simulation benchmark. Errors in generated CoT (VLM hallucinations, latent-action noise) propagate directly into action prediction. CoT adds no value when the task is direct perception-to-action (pick-and-place) without multi-step planning.

VLM backbone preservation. Robot-only training erodes PaliGemma2's MMBench/GQA scores; balanced co-training restores them β€” a side-finding that justifies VL co-training even when robot-task gain alone might not.

Significance

This is the first controlled ablation of co-training modalities Γ— training schedules at LBM scale, and it overturns three pieces of folk wisdom that had been propagating through 2025–2026:

  1. "More data modalities are always better." Wrong β€” discrete-action token modalities (FAST, VQ-VAE) and latent video tokens add noise without benefit at scale.
  2. "Chain-of-thought is universally helpful." Wrong for direct perception-to-action manipulation β€” CoT only helps when there is real multi-step planning to expose. This corroborates from a different angle the Hybrid Training / Actions as Language line that ECoT's gain comes from training-time supervision, not inference-time conditioning.
  3. "Cross-embodiment data should be co-trained throughout." Wrong β€” it pays off when used in phase 1 only, then dropped for phase-2 specialization. The opposite of what naΓ―ve "more data is more better" intuition predicts.

It also gives the community three concrete recipes:

  • For VL-like modalities (web VL + human-video captions): two-phase full co-training.
  • For cross-embodiment robot data: two-phase first-phase only.
  • Skip discrete action tokenization unless you are explicitly in the low-data regime.

This positions the paper as the empirical companion to Knowledge Insulation (which formalizes how to protect the VLM during action training) β€” KI says how to avoid corrupting the backbone, this paper says what to feed alongside the robot data. Together they form the most complete public training-recipe answer for an LBM in 2026.

It also disambiguates the VLM4VLA result: VLM benchmark score doesn't predict VLA success, but co-training on VL data is what actually keeps the VLM from forgetting what it knew β€” which is why the right choice of co-training data (this paper) and the right gradient routing (KI) end up mattering more than the absolute VLM benchmark number (VLM4VLA).

Links

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️