ICLR 2026 LBM Cotraining - Heungwoo/research GitHub Wiki
Venue: arXiv preprint (Feb 1, 2026) β Toyota Research Institute (TRI) arXiv: 2602.01067 Category: Training Approach / Data Strategy (empirical study) Trend tag: Trend 2 (training recipes) Β· Trend 4 (data scaling) Scale of study: 4,000 hours of robot/human manipulation + 50M VL samples Β· 89 policies trained Β· 58,000 sim rollouts + 2,835 real-robot rollouts
flowchart LR
subgraph Mod[5 co-training modalities]
M1[1. Standard VL data<br/>RoboPoint + RefSpatial<br/>50M samples]
M2[2. Dense robot annotations<br/>scripted + GPT-5 captions<br/>523h]
M3[3. Cross-embodiment robot<br/>OXE-Ramen 1,150h<br/>12 robots / 924 tasks]
M4[4. Human videos<br/>2,271h Ego4D / EgoDex / SSv2<br/>latent actions OR captions]
M5[5. Discrete robot tokens<br/>FAST or VQ-VAE<br/>523h]
end
subgraph LBM[VLA / LBM under test]
VLM[PaliGemma2-3B-PT<br/>vision-language backbone]
AH[ActionFT<br/>8-layer flow-matching transformer]
VLM -- adaLN conditioning --> AH
end
subgraph Strat[3 training strategies]
S1[Single-phase joint]
S2[2-phase: 1st-phase only]
S3[2-phase: full co-training]
end
Mod --> LBM
Strat --> LBM
LBM --> Eval[Sim 21 tasks Β· Real 9+ tasks<br/>distribution shift / unseen / language following]
Large Behavior Models (LBMs) β multi-task imitation policies that scale to dexterous manipulation β are bottlenecked by insufficient robot data coverage. The community's response is co-training with heterogeneous data (web VL, OXE cross-embodiment, human videos, action tokens, β¦), but until now the field has had no controlled comparison of which modalities and which training schedules actually pay off.
Concretely the paper attacks four open questions:
- Which co-training data modalities help an LBM, and how much?
- Should the modality be used in the first phase only (pretraining), throughout (joint co-training), or in a single-phase mix?
- Does chain-of-thought conditioning on co-training-derived traces improve action prediction?
- Are gains additive when modalities are stacked?
Backbone (LBM under study). PaliGemma2-3B-PT VLM (google/paligemma2-3b-pt-224, fine-tuned during co-training β its trainability is exactly what makes the MMBench/GQA preservation finding possible) feeds an 8-layer ActionFT flow-matching transformer. A single global conditioning token is extracted from the VLM's last four layers and injected into every ActionFT layer via an adaLN MLP (alongside the flow-matching timestep). Action chunk = 16 steps (~1.6 s at 10 Hz) of relative end-effector poses w.r.t. the station base frame + gripper widths. A small auxiliary discrete-token head is added when training with discrete-action losses (CE) alongside flow matching (FM).
Five co-training modalities studied:
| # | Modality | Source | Format |
|---|---|---|---|
| 1 | Standard VL | RoboPoint (1.3M / 8.2M QA) + RefSpatial (2.5M / 20M QA) | VQA, object localization, spatial reasoning |
| 2 | Dense robot annotations | TRI-Ramen (β523h) with two caption types: (a) scripted-heuristic motion descriptions, (b) GPT-5 generated rich descriptions | Per-frame language captions at ~1β2 s intervals |
| 3 | Cross-embodiment robot | OXE-Ramen (curated Open X-Embodiment) β 12 setups, 924 tasks, 466k demos, 1,150h | Native continuous actions, diverse morphologies |
| 4 | Human videos | Ego4D, EgoDex, Something-Something V2, Epic Kitchen, HoloAssist β 2,271h filtered. Two representations: (a) latent action tokens via Latent Action Model (codebook 32, 8 tokens/chunk), (b) GPT-5 motion captions | Either token sequence (CE loss) or caption (CE) |
| 5 | Discrete robot tokens | TRI-Ramen retokenized two ways: (a) FAST (~42.1 tokens / chunk, vocab 2,048), (b) VQ-VAE (8 tokens, codebook 32) | Compressed action tokens (CE) |
Three training strategies:
- Single-phase joint β target robot data + co-training modalities mixed in one phase.
- Two-phase, 1st-phase-only β co-training data in phase 1, then target robot continuous actions in phase 2.
- Two-phase, full co-training β co-training data in phase 1, then target + co-training jointly in phase 2.
Loss. Flow-matching for continuous actions, cross-entropy for any discrete token / language stream, weighted sum with computed masks.
Chain-of-thought conditioning experiment. Three CoT variants tested on three trace types (scripted captions, VLM captions, latent actions): (i) train with CoT 50% of the time, infer without CoT; (ii) same training, generate CoT then condition actions on it at inference; (iii) always condition on CoT during training and inference.
Simulation β 21 tasks (13 seen + 8 unseen), 50 rollouts each, 58k+ total rollouts.
| Setting | Unseen-task success |
|---|---|
| No co-training (target only) | β 36.2% |
| All effective modalities stacked | 72.6% (+36.4 pp) |
Distribution-shift robustness (lighting, textures, camera, backgrounds) climbs in lockstep with each effective modality.
Real-world (dual Franka Panda, 2,835 rollouts).
| Eval | Baseline | Final stacked LBM |
|---|---|---|
| Language following (45 / cond.) | β 24.1% | 69.4% (+45.3 pp) |
| Long-horizon dex tasks fine-tune (200 demos, 30 rollouts) | 67.4% (FT'd baseline) / 47.3% (single-task untrained) | 90.2% |
Per-modality ranking (effective β ineffective):
- β Standard VL data β largest single-modality gain
- β VLM-generated robot annotations β > scripted heuristics
- β Cross-embodiment robot data β most useful in phase-1 only
- β VLM-generated human-video annotations β sustained co-training helps
- β FAST discrete tokens β no significant benefit; degrades unseen tasks
- β Latent action tokens from human video β only in low-data regime, diminishing
- β VQ-VAE robot tokens β marginal/null
Chain-of-thought conditioning. All three CoT variants either match or degrade vs. implicit two-phase co-training on the simulation benchmark. Errors in generated CoT (VLM hallucinations, latent-action noise) propagate directly into action prediction. CoT adds no value when the task is direct perception-to-action (pick-and-place) without multi-step planning.
VLM backbone preservation. Robot-only training erodes PaliGemma2's MMBench/GQA scores; balanced co-training restores them β a side-finding that justifies VL co-training even when robot-task gain alone might not.
This is the first controlled ablation of co-training modalities Γ training schedules at LBM scale, and it overturns three pieces of folk wisdom that had been propagating through 2025β2026:
- "More data modalities are always better." Wrong β discrete-action token modalities (FAST, VQ-VAE) and latent video tokens add noise without benefit at scale.
- "Chain-of-thought is universally helpful." Wrong for direct perception-to-action manipulation β CoT only helps when there is real multi-step planning to expose. This corroborates from a different angle the Hybrid Training / Actions as Language line that ECoT's gain comes from training-time supervision, not inference-time conditioning.
- "Cross-embodiment data should be co-trained throughout." Wrong β it pays off when used in phase 1 only, then dropped for phase-2 specialization. The opposite of what naΓ―ve "more data is more better" intuition predicts.
It also gives the community three concrete recipes:
- For VL-like modalities (web VL + human-video captions): two-phase full co-training.
- For cross-embodiment robot data: two-phase first-phase only.
- Skip discrete action tokenization unless you are explicitly in the low-data regime.
This positions the paper as the empirical companion to Knowledge Insulation (which formalizes how to protect the VLM during action training) β KI says how to avoid corrupting the backbone, this paper says what to feed alongside the robot data. Together they form the most complete public training-recipe answer for an LBM in 2026.
It also disambiguates the VLM4VLA result: VLM benchmark score doesn't predict VLA success, but co-training on VL data is what actually keeps the VLM from forgetting what it knew β which is why the right choice of co-training data (this paper) and the right gradient routing (KI) end up mattering more than the absolute VLM benchmark number (VLM4VLA).
- arXiv: 2602.01067 (Feb 1, 2026)
- HTML: https://arxiv.org/html/2602.01067
- Companion TRI paper: A Careful Examination of Large Behavior Models (LBM definitions, dataset framing)
- In-depth review: LBM Co-training Study (long-form)
- Knowledge Insulation β how to protect the VLM during action training
- Hybrid Training β ECoT curriculum that drops reasoning at inference
- Actions as Language β language relabeling of subtasks
- VLM4VLA β VLM-benchmark-vs-VLA-perf disambiguation
- Human-Video Pretraining β adjacent on human-video co-training
- FASTER β discrete-token tokenizer this paper finds unhelpful at LBM scale
- Survey: VLA & Manipulation
β Back to ICLR-2026