RSS 2026 LDA 1B - Heungwoo/research GitHub Wiki

LDA-1B โ€” Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestion

Venue: RSS 2026 (Imitation Learning session) ยท Authors: Jiangran Lyu, Kai Liu, Xuheng Zhang, Haoran Liao, โ€ฆ Ming-Yu Liu, Zhizheng Zhang, et al. (PKU / Galbot / NVIDIA lineage) ยท arXiv: 2602.12215 Category: World-model-based robot foundation model Trend tag: RSS 2026 thread 3 โ€” video/world models vs the VLA backbone

Compiled from the verified RSS 2026 abstract; arXiv ID cross-confirmed via Qwen-RobotManip's reference list.

Key figure

LDA-1B overview (Figure 1 of arXiv 2602.12215, ยฉ the authors)

Figure 1 of the LDA-1B paper. Left: the three-tier data taxonomy of EI-30k โ€” high-quality action data (real/sim robot + human), suboptimal/noisy action data, and actionless human videos, 30k+ hours in a unified LeRobot-style format with aligned coordinate systems. Center: each tier feeds a different objective โ€” all tasks for high-quality data, dynamics for noisy data, forecasting for actionless video โ€” into one model (the figure states 1.6B parameters, ~10ร— prior UWM instantiations), with visual forecasting in DINO latent space rather than pixels. Right: headline real-robot comparisons vs ฯ€0.5 โ€” contact-rich 42โ†’63, dexterous 37โ†’85, long-horizon 27โ†’50 โ€” and the data-efficient fine-tuning result (adding non-expert data: ฯ€0.5 drops 55โ†’40 while LDA-1B rises 60โ†’70).

Problem

Behavior-cloning robot foundation models imitate expert actions but discard the transferable dynamics knowledge embedded in heterogeneous embodied data โ€” failures, suboptimal rollouts, action-free video. The Unified World Model (UWM) formulation could exploit all of it, but prior instantiations don't scale to foundation level: coarse data usage, fragmented datasets, pixel-space redundancy.

Method

  • Joint objectives: dynamics learning + policy learning + visual forecasting in one model, with distinct roles assigned to data of different quality tiers (expert data supervises the policy; low-quality data still teaches dynamics).
  • EI-30k: a standardized embodied-interaction corpus of >30,000 hours of human and robot trajectories in a unified format.
  • Structured DINO latent space for dynamics โ€” avoids modeling pixel-level appearance, which is what previously blocked scaling.
  • Mixed-frequency multi-modal diffusion transformer handling asynchronous vision and action streams; stable at 1B parameters.

Results (as reported)

  • Outperforms prior methods (incl. ฯ€0.5) by up to +21% (contact-rich), +48% (dexterous), +23% (long-horizon) in simulation and real world.
  • Data-efficient fine-tuning: +10% by ingesting the 30% of low-quality trajectories that BC pipelines normally discard as harmful.
  • Code and data release stated.

Significance

The strongest scaling result yet for the "learn dynamics, not just actions" camp ([Review-World-Models]]) โ€” and the direct empirical counter to pure behavior cloning at foundation scale. The low-quality-data result operationalizes what LBM-style curation studies only hint at: bad trajectories are dynamics data. Cited as a representation-alignment reference by [Qwen-RobotManip ("LDA-1B... universal embodied data ingestion"); sits alongside mimic-video as RSS 2026's two-pronged challenge to the VLM-backbone monopoly.

โ† RSS 2026 survey ยท Home