RSS 2026 LDA 1B - Heungwoo/research GitHub Wiki
LDA-1B โ Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestion
Venue: RSS 2026 (Imitation Learning session) ยท Authors: Jiangran Lyu, Kai Liu, Xuheng Zhang, Haoran Liao, โฆ Ming-Yu Liu, Zhizheng Zhang, et al. (PKU / Galbot / NVIDIA lineage) ยท arXiv: 2602.12215 Category: World-model-based robot foundation model Trend tag: RSS 2026 thread 3 โ video/world models vs the VLA backbone
Compiled from the verified RSS 2026 abstract; arXiv ID cross-confirmed via Qwen-RobotManip's reference list.
Key figure

Figure 1 of the LDA-1B paper. Left: the three-tier data taxonomy of EI-30k โ high-quality action data (real/sim robot + human), suboptimal/noisy action data, and actionless human videos, 30k+ hours in a unified LeRobot-style format with aligned coordinate systems. Center: each tier feeds a different objective โ all tasks for high-quality data, dynamics for noisy data, forecasting for actionless video โ into one model (the figure states 1.6B parameters, ~10ร prior UWM instantiations), with visual forecasting in DINO latent space rather than pixels. Right: headline real-robot comparisons vs ฯ0.5 โ contact-rich 42โ63, dexterous 37โ85, long-horizon 27โ50 โ and the data-efficient fine-tuning result (adding non-expert data: ฯ0.5 drops 55โ40 while LDA-1B rises 60โ70).
Problem
Behavior-cloning robot foundation models imitate expert actions but discard the transferable dynamics knowledge embedded in heterogeneous embodied data โ failures, suboptimal rollouts, action-free video. The Unified World Model (UWM) formulation could exploit all of it, but prior instantiations don't scale to foundation level: coarse data usage, fragmented datasets, pixel-space redundancy.
Method
- Joint objectives: dynamics learning + policy learning + visual forecasting in one model, with distinct roles assigned to data of different quality tiers (expert data supervises the policy; low-quality data still teaches dynamics).
- EI-30k: a standardized embodied-interaction corpus of >30,000 hours of human and robot trajectories in a unified format.
- Structured DINO latent space for dynamics โ avoids modeling pixel-level appearance, which is what previously blocked scaling.
- Mixed-frequency multi-modal diffusion transformer handling asynchronous vision and action streams; stable at 1B parameters.
Results (as reported)
- Outperforms prior methods (incl. ฯ0.5) by up to +21% (contact-rich), +48% (dexterous), +23% (long-horizon) in simulation and real world.
- Data-efficient fine-tuning: +10% by ingesting the 30% of low-quality trajectories that BC pipelines normally discard as harmful.
- Code and data release stated.
Significance
The strongest scaling result yet for the "learn dynamics, not just actions" camp ([Review-World-Models]]) โ and the direct empirical counter to pure behavior cloning at foundation scale. The low-quality-data result operationalizes what LBM-style curation studies only hint at: bad trajectories are dynamics data. Cited as a representation-alignment reference by [Qwen-RobotManip ("LDA-1B... universal embodied data ingestion"); sits alongside mimic-video as RSS 2026's two-pronged challenge to the VLM-backbone monopoly.
โ RSS 2026 survey ยท Home