ICLR 2026 Disentangled FwdInv - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 (Poster) ยท Authors: Wenyao Zhang, Bozhou Zhang, Zekun Qi, Wenjun Zeng, Xin Jin, Li Zhang ยท arXiv:2604.16391 ยท Category: Data / memory / representation for manipulation ยท Trend tag: Representation / video pretraining
Official title: Disentangled Robot Learning via Separate Forward and Inverse Dynamics Pretraining.
flowchart LR
Vid[Human + robot videos<br/>action-free web video] --> GFDM[GFDM<br/>General Forward Dynamics Model<br/>future-frame prediction]
Vid --> GIDM[GIDM<br/>General Inverse Dynamics Model<br/>self-supervised latent actions]
GFDM --> Unify[Unified architecture]
GIDM --> Unify
Unify --> FT[End-to-end fine-tuning<br/>on downstream tasks]
FT --> Act[Action prediction]
VLA models suffer a misalignment between 2D image forecasting and 3D action prediction, and vision-action entangled training prevents them from learning from large-scale, action-free web video. Forecasting pixels and predicting actions pull the model in different directions when trained jointly.
DeFI decouples visual forward and inverse dynamics pretraining so each exploits its own best data source, disentangling video generation from action prediction:
- GFDM (General Forward Dynamics Model) โ pretrained on diverse human and robot videos for future prediction.
- GIDM (General Inverse Dynamics Model) โ trained via self-supervised learning to infer latent actions from unlabeled video transitions, unlocking action-free video.
The two are then integrated into a unified architecture for end-to-end downstream fine-tuning.
Evaluated on CALVIN ABC-D and SimplerEnv-Fractal. Reported figures: CALVIN ABC-D average task length 4.51; SimplerEnv-Fractal 51.2% success; real-world deployment 81.3% success โ described as significantly outperforming prior methods, with gains in efficiency, scalability, and generalization.
Reframes VLA pretraining as two separable problems โ "what happens next" (forward) and "what action caused this transition" (inverse) โ each matched to the data it needs. Letting GIDM mine action-free video for latent actions is a scalable route around the perennial scarcity of action-labeled robot data.
โ Back to ICLR-2026