ICLR 2026 Disentangled FwdInv - Heungwoo/research GitHub Wiki

DeFI โ€” disentangled forward & inverse dynamics pretraining

Venue: ICLR 2026 (Poster) ยท Authors: Wenyao Zhang, Bozhou Zhang, Zekun Qi, Wenjun Zeng, Xin Jin, Li Zhang ยท arXiv:2604.16391 ยท Category: Data / memory / representation for manipulation ยท Trend tag: Representation / video pretraining

Official title: Disentangled Robot Learning via Separate Forward and Inverse Dynamics Pretraining.

Approach diagram

flowchart LR
  Vid[Human + robot videos<br/>action-free web video] --> GFDM[GFDM<br/>General Forward Dynamics Model<br/>future-frame prediction]
  Vid --> GIDM[GIDM<br/>General Inverse Dynamics Model<br/>self-supervised latent actions]
  GFDM --> Unify[Unified architecture]
  GIDM --> Unify
  Unify --> FT[End-to-end fine-tuning<br/>on downstream tasks]
  FT --> Act[Action prediction]
Loading

Problem

VLA models suffer a misalignment between 2D image forecasting and 3D action prediction, and vision-action entangled training prevents them from learning from large-scale, action-free web video. Forecasting pixels and predicting actions pull the model in different directions when trained jointly.

Method

DeFI decouples visual forward and inverse dynamics pretraining so each exploits its own best data source, disentangling video generation from action prediction:

  • GFDM (General Forward Dynamics Model) โ€” pretrained on diverse human and robot videos for future prediction.
  • GIDM (General Inverse Dynamics Model) โ€” trained via self-supervised learning to infer latent actions from unlabeled video transitions, unlocking action-free video.

The two are then integrated into a unified architecture for end-to-end downstream fine-tuning.

Results

Evaluated on CALVIN ABC-D and SimplerEnv-Fractal. Reported figures: CALVIN ABC-D average task length 4.51; SimplerEnv-Fractal 51.2% success; real-world deployment 81.3% success โ€” described as significantly outperforming prior methods, with gains in efficiency, scalability, and generalization.

Significance

Reframes VLA pretraining as two separable problems โ€” "what happens next" (forward) and "what action caused this transition" (inverse) โ€” each matched to the data it needs. Letting GIDM mine action-free video for latent actions is a scalable route around the perennial scarcity of action-labeled robot data.

Links

Related pages

โ† Back to ICLR-2026

โš ๏ธ **GitHub.com Fallback** โš ๏ธ