CVPR 2026 DiT4DiT - Heungwoo/research GitHub Wiki

DiT4DiT โ€” Jointly Modeling Video Dynamics and Actions for Generalizable Robot Control

Venue: CVPR 2026 Authors / affiliations: Teli Ma, Jia Zheng, Zifan Wang, Chunli Jiang, Andy Cui, Junwei Liang*, Shuo Yang* โ€” Mondo Robotics, HKUST(GZ), HKUST Category: VLA Architecture (cascaded video+action DiT) Trend tag: Trend 3 (world-model + RL)

Approach diagram

flowchart LR
  OBS["current obs"] --> VIDEO_DIT["Video-DiT (Cosmos-Predict2.5-2B)<br/>flow-match video"]
  VIDEO_DIT --> FEAT["intermediate denoising<br/>hidden-state features (forward hook @ ฯ„_f)"]
  FEAT --> ACT_DIT["Action-DiT<br/>flow-match action chunks"]
  OBS --> ACT_DIT
  ACT_DIT --> ACT["action chunk"]
Loading

Problem

Most VLAs predict actions directly; world-model VLAs (Cosmos Policy, VideoVLA) predict frames and actions jointly. The latter helps generalization but adds a heavy generator on the critical path โ€” and conditioning the policy on fully reconstructed future frames forces the expensive full denoising trajectory at inference and ties the policy to pixel-level details. The question DiT4DiT asks: can a video diffusion model help action prediction without reconstructing complete future frames?

Method

A cascaded video-DiT + action-DiT trained end-to-end with a dual flow-matching objective. Crucially, the action-DiT is not conditioned on reconstructed future frames. Instead, a forward-hook mechanism intercepts intermediate denoising hidden-state activations of the video-DiT at a fixed flow timestep ฯ„_f and feeds them to the action-DiT via cross-attention as "physics-aware visual tokens" โ€” no full video reconstruction required.

The video backbone is Cosmos-Predict2.5-2B (NVIDIA, foundation-scale), and the full model has ~2.2B trainable parameters. Training uses an asymmetric tri-timestep scheme with three independent timesteps: ฯ„_v (video, uniform on [0,1]), ฯ„_f (feature extraction, fixed/deterministic for stable representations), and ฯ„_a (action, Beta-sampled to focus on critical control phases). Decoupling these lets the action module learn inverse dynamics without joint-optimization conflicts.

Results

  • LIBERO: 98.6 % average across the four suites (Spatial 98.4 / Object 99.6 / Goal 98.6 / Long 97.6) โ€” among the strongest reported, ahead of ฯ€0.5 (96.9), OpenVLA-OFT (97.1), CogVLA (97.4).
  • RoboCasa-GR1: 50.8 % average over 24 tasks (vs. GR00T-N1.5 41.8) โ€” using only ~15 % of the pre-training data (241,450 GR1 episodes) of comparable baselines.
  • Real-world Unitree G1: outperforms GR00T-N1.5 across seven household tasks (e.g. Arrange Flower 75 % vs 25 %, Stack Cup 60 % vs 25 %), with strong zero-shot generalization to unseen objects. A parameter-matched Qwen3-DiT baseline collapsed to ~0 % in the real world.

Note: arXiv abstract/HTML reports RoboCasa-GR1 = 50.8 % (vs GR00T-N1.5 41.8); the project page lists a different figure (56.7 % vs GR00T-N1.6 47.8). The arXiv value is used here.

Decisive ablation: extracting features at 1 denoising step is best, with success rate degrading monotonically as steps increase โ€” more denoising forces the hidden states to over-commit to pixel-level details of a specific reconstructed future. This validates the core claim that high-frequency control does not need complete frame reconstruction.

Significance

DiT4DiT shows the world-model signal can be tapped without reconstructing future frames: intermediate denoising features from a foundation-scale video DiT (Cosmos-Predict2.5-2B), extracted at a single early flow step, are a stronger and cheaper action condition than the final generated pixels. This sidesteps the cost and pixel-overfitting of frame-conditioned world-model VLAs while still inheriting a large pretrained video prior. Sits adjacent to ฯ€0.7's BAGEL-14B (huge separate world model) and Unified Diffusion VLA (one network, co-denoised). DiT4DiT's distinct position: a cascaded design that conditions actions on the video model's internal denoising state rather than its output.

Links

Related pages

โ† Back to CVPR-2026

โš ๏ธ **GitHub.com Fallback** โš ๏ธ