RSS 2026 LBM Cotraining Study - Heungwoo/research GitHub Wiki
A Systematic Study of Data Modalities and Strategies for Co-training Large Behavior Models
Venue: RSS 2026 (Manipulation session) ยท Authors: Fanqi Lin, Kushal Arora, Jean Mercat, Haruki Nishimura, Paarth Shah, โฆ Jose Barreiros โ Toyota Research Institute ยท arXiv: 2602.01067 ยท project Category: Empirical co-training study (LBM series) Trend tag: RSS 2026 cross-cutting anchor โ data recipes
Compiled from the verified RSS 2026 abstract and the paper's Fig. 1.
Key figure

Figure 1 of the study. Top: the experimental matrix โ target robot data (sim + real) against the five co-training modalities: standard VL data (VQA/captioning/detection), language annotations for robot data (episode-, scripted-frame-, and VLM-frame-level), cross-embodiment robot data, human videos, and discrete robot action tokens (FAST + VQ-VAE). Bottom-left: the fixed policy under study โ a VLM backbone + Action Flow Transformer, whose text head consumes annotations/captions/action-tokens while the flow head produces continuous actions. Bottom-right: the evaluation grid โ sim seen/distribution-shift/unseen tasks plus real-world language-following and long-horizon dexterous tasks. Holding the architecture fixed while varying only the data is what makes the 89-policy comparison clean.
Problem
LBM generalization is capped by robot-data coverage; co-training with heterogeneous data is the standard workaround โ but which modalities and which training strategies actually help has never been measured systematically at scale.
Method
- Five co-training modalities tested: (1) standard vision-language data, (2) dense language annotations on robot trajectories, (3) cross-embodiment robot data, (4) human videos, (5) discrete robot action tokens โ across single- and multi-phase training strategies.
- Scale: 4,000 h of robot + human manipulation data, 50M VL samples, 89 trained VLA policies, evaluated over 58,000 simulation rollouts and 2,835 real-world rollouts.
Results (as reported)
- VL data and cross-embodiment robot data substantially improve generalization to distribution shifts, unseen tasks, and language following.
- Discrete action-token variants yield no statistically significant benefit โ replicating the LBM-1 finding at much larger scale.
- Effective modalities combine cumulatively, and the combined recipe enables rapid fine-tuning to unseen long-horizon dexterous tasks.
Significance
The definitive sequel to the ICLR LBM co-training study โ TRI turning its co-training findings into the field's reference experiment (89 policies is an order of magnitude beyond typical ablation budgets). Three results now stand as near-consensus: (a) VL co-training is generalization infrastructure (echoed by Qwen-RobotManip's +8.2 pp OOD result and the Qwen program's ฮป-weighted streams), (b) cross-embodiment data transfers (the Cross-Embodiment agenda), and (c) discrete action tokens don't help โ now triply replicated (LBM-1, this study, and the Qwen program's wholesale omission). Also corroborates human video as a useful modality, bridging to H2R emergence.
โ RSS 2026 survey ยท Home