CVPR 2026 RDT2 - Heungwoo/research GitHub Wiki
Venue: CVPR 2026 Category: VLA Architecture (Continuous Diffusion / scaling) Trend tag: Trend 1 Affiliations: Tsinghua THU-ML
flowchart LR
UMI["UMI demos<br/>10k+ hours, hand-held gripper"] --> RDT2["RDT2 โ Qwen2.5-VL-7B-Instruct<br/>(2 wrist fisheye imgs + language)"]
RDT2 --> S1["Stage 1: RDT2-VQ<br/>RVQ tokens + cross-entropy"]
RDT2 --> S2["Stage 2: RDT2-FM<br/>400M RDT flow-matching expert"]
S2 --> S3["Stage 3: RDT2-UltraFast<br/>one-step distillation"]
S1 --> ACT["action chunk"]
S3 --> ACT
ROBOT_NEW["unseen robot at deploy"] --> RDT2
RDT-1B (1.2 B) showed continuous diffusion VLAs can scale across embodiments via a unified action space. The question RDT2 attacks: how far does this scale, and what is the data ceiling? Real-robot teleop data is bottlenecked at hundreds of hours; UMI (universal manipulation interface, hand-held gripper) data scales to 10 000+ hours.
- Build on a Qwen2.5-VL-7B-Instruct VLM backbone, taking two wrist-view fisheye images + a language instruction as input.
- Train on 10 000+ hours of UMI demonstrations (one of the largest open-source robot datasets) collected with an enhanced, embodiment-agnostic UMI.
- A three-stage recipe that aligns the VLM's discrete linguistic knowledge with continuous control:
- Stage 1 โ RDT2-VQ: discretize action chunks with RVQ (Residual Vector Quantization) and train the VLM with cross-entropy. A 0.8 s chunk @ 30 Hz compresses to 27 tokens (โ1/3 of FAST, 1/8 of binning); requires 27 autoregressive passes per chunk. This discrete head also enables RL.
- Stage 2 โ RDT2-FM: attach a 400 M RDT flow-matching action expert (improved RDT model) for continuous action generation without quantization error โ one Qwen pass + ~5 expert passes. RDT2-FM-Post adds real-robot data (UR, Franka).
- Stage 3 โ RDT2-UltraFast: distill RDT2-FM into a one-step diffusion policy โ one Qwen pass + one expert pass for real-time inference.
- Evaluate zero-shot transfer to robots not seen in training.
Among the first models to simultaneously zero-shot generalize across the "4U" axes โ unseen embodiment, scene, object, and language โ since UMI demonstrations are robot-agnostic in their hand-held form. Reported to outperform SOTA baselines such as ฯ0-FAST and ฯ0.5 on dexterous, long-horizon, and dynamic downstream tasks. Demos include table tennis (predictive ball-trajectory planning under a 1 m/s arm limit), archery (~100 ms reaction), and fabric folding generalizing to unseen textures/sizes. The public materials emphasize qualitative demos; per-task quantitative tables were not extractable from the project page / arXiv HTML at audit time.
This is the cleanest scaling-law evidence for the UMI-data โ general policy pathway. UMI's bet is that hand-held demonstrations are an underutilized data source 10โ100ร larger than teleop; RDT2 validates that this is enough to learn a generalist policy with zero-shot transfer. Most relevant predecessors: DexUMI (UMI for dexterous hands), X-VLA (soft-prompt cross-embodiment).
- arXiv: 2602.03310
- VLA Architecture review ยงC (continuous diffusion)
- Cross-Embodiment review
- DexUMI ยท X-VLA
- CVPR 2026 survey
โ Back to CVPR-2026