RSS 2026 LBM Cotraining Study - Heungwoo/research GitHub Wiki

A Systematic Study of Data Modalities and Strategies for Co-training Large Behavior Models

Venue: RSS 2026 (Manipulation session) ยท Authors: Fanqi Lin, Kushal Arora, Jean Mercat, Haruki Nishimura, Paarth Shah, โ€ฆ Jose Barreiros โ€” Toyota Research Institute ยท arXiv: 2602.01067 ยท project Category: Empirical co-training study (LBM series) Trend tag: RSS 2026 cross-cutting anchor โ€” data recipes

Compiled from the verified RSS 2026 abstract and the paper's Fig. 1.

Key figure

Study overview (Figure 1 of arXiv 2602.01067, ยฉ Toyota Research Institute)

Figure 1 of the study. Top: the experimental matrix โ€” target robot data (sim + real) against the five co-training modalities: standard VL data (VQA/captioning/detection), language annotations for robot data (episode-, scripted-frame-, and VLM-frame-level), cross-embodiment robot data, human videos, and discrete robot action tokens (FAST + VQ-VAE). Bottom-left: the fixed policy under study โ€” a VLM backbone + Action Flow Transformer, whose text head consumes annotations/captions/action-tokens while the flow head produces continuous actions. Bottom-right: the evaluation grid โ€” sim seen/distribution-shift/unseen tasks plus real-world language-following and long-horizon dexterous tasks. Holding the architecture fixed while varying only the data is what makes the 89-policy comparison clean.

Problem

LBM generalization is capped by robot-data coverage; co-training with heterogeneous data is the standard workaround โ€” but which modalities and which training strategies actually help has never been measured systematically at scale.

Method

  • Five co-training modalities tested: (1) standard vision-language data, (2) dense language annotations on robot trajectories, (3) cross-embodiment robot data, (4) human videos, (5) discrete robot action tokens โ€” across single- and multi-phase training strategies.
  • Scale: 4,000 h of robot + human manipulation data, 50M VL samples, 89 trained VLA policies, evaluated over 58,000 simulation rollouts and 2,835 real-world rollouts.

Results (as reported)

  • VL data and cross-embodiment robot data substantially improve generalization to distribution shifts, unseen tasks, and language following.
  • Discrete action-token variants yield no statistically significant benefit โ€” replicating the LBM-1 finding at much larger scale.
  • Effective modalities combine cumulatively, and the combined recipe enables rapid fine-tuning to unseen long-horizon dexterous tasks.

Significance

The definitive sequel to the ICLR LBM co-training study โ€” TRI turning its co-training findings into the field's reference experiment (89 policies is an order of magnitude beyond typical ablation budgets). Three results now stand as near-consensus: (a) VL co-training is generalization infrastructure (echoed by Qwen-RobotManip's +8.2 pp OOD result and the Qwen program's ฮป-weighted streams), (b) cross-embodiment data transfers (the Cross-Embodiment agenda), and (c) discrete action tokens don't help โ€” now triply replicated (LBM-1, this study, and the Qwen program's wholesale omission). Also corroborates human video as a useful modality, bridging to H2R emergence.

โ† RSS 2026 survey ยท Home