CVPR 2026 DiT4DiT - Heungwoo/research GitHub Wiki
Venue: CVPR 2026 Authors / affiliations: Teli Ma, Jia Zheng, Zifan Wang, Chunli Jiang, Andy Cui, Junwei Liang*, Shuo Yang* โ Mondo Robotics, HKUST(GZ), HKUST Category: VLA Architecture (cascaded video+action DiT) Trend tag: Trend 3 (world-model + RL)
flowchart LR
OBS["current obs"] --> VIDEO_DIT["Video-DiT (Cosmos-Predict2.5-2B)<br/>flow-match video"]
VIDEO_DIT --> FEAT["intermediate denoising<br/>hidden-state features (forward hook @ ฯ_f)"]
FEAT --> ACT_DIT["Action-DiT<br/>flow-match action chunks"]
OBS --> ACT_DIT
ACT_DIT --> ACT["action chunk"]
Most VLAs predict actions directly; world-model VLAs (Cosmos Policy, VideoVLA) predict frames and actions jointly. The latter helps generalization but adds a heavy generator on the critical path โ and conditioning the policy on fully reconstructed future frames forces the expensive full denoising trajectory at inference and ties the policy to pixel-level details. The question DiT4DiT asks: can a video diffusion model help action prediction without reconstructing complete future frames?
A cascaded video-DiT + action-DiT trained end-to-end with a dual flow-matching objective. Crucially, the action-DiT is not conditioned on reconstructed future frames. Instead, a forward-hook mechanism intercepts intermediate denoising hidden-state activations of the video-DiT at a fixed flow timestep ฯ_f and feeds them to the action-DiT via cross-attention as "physics-aware visual tokens" โ no full video reconstruction required.
The video backbone is Cosmos-Predict2.5-2B (NVIDIA, foundation-scale), and the full model has ~2.2B trainable parameters. Training uses an asymmetric tri-timestep scheme with three independent timesteps: ฯ_v (video, uniform on [0,1]), ฯ_f (feature extraction, fixed/deterministic for stable representations), and ฯ_a (action, Beta-sampled to focus on critical control phases). Decoupling these lets the action module learn inverse dynamics without joint-optimization conflicts.
- LIBERO: 98.6 % average across the four suites (Spatial 98.4 / Object 99.6 / Goal 98.6 / Long 97.6) โ among the strongest reported, ahead of ฯ0.5 (96.9), OpenVLA-OFT (97.1), CogVLA (97.4).
- RoboCasa-GR1: 50.8 % average over 24 tasks (vs. GR00T-N1.5 41.8) โ using only ~15 % of the pre-training data (241,450 GR1 episodes) of comparable baselines.
- Real-world Unitree G1: outperforms GR00T-N1.5 across seven household tasks (e.g. Arrange Flower 75 % vs 25 %, Stack Cup 60 % vs 25 %), with strong zero-shot generalization to unseen objects. A parameter-matched Qwen3-DiT baseline collapsed to ~0 % in the real world.
Note: arXiv abstract/HTML reports RoboCasa-GR1 = 50.8 % (vs GR00T-N1.5 41.8); the project page lists a different figure (56.7 % vs GR00T-N1.6 47.8). The arXiv value is used here.
Decisive ablation: extracting features at 1 denoising step is best, with success rate degrading monotonically as steps increase โ more denoising forces the hidden states to over-commit to pixel-level details of a specific reconstructed future. This validates the core claim that high-frequency control does not need complete frame reconstruction.
DiT4DiT shows the world-model signal can be tapped without reconstructing future frames: intermediate denoising features from a foundation-scale video DiT (Cosmos-Predict2.5-2B), extracted at a single early flow step, are a stronger and cheaper action condition than the final generated pixels. This sidesteps the cost and pixel-overfitting of frame-conditioned world-model VLAs while still inheriting a large pretrained video prior. Sits adjacent to ฯ0.7's BAGEL-14B (huge separate world model) and Unified Diffusion VLA (one network, co-denoised). DiT4DiT's distinct position: a cascaded design that conditions actions on the video model's internal denoising state rather than its output.
- arXiv: 2603.10448
โ Back to CVPR-2026