CVPR 2026 RDT2 - Heungwoo/research GitHub Wiki

RDT2 โ€” Exploring the Scaling Limit of UMI Data Toward Zero-Shot Cross-Embodiment Generalization

Venue: CVPR 2026 Category: VLA Architecture (Continuous Diffusion / scaling) Trend tag: Trend 1 Affiliations: Tsinghua THU-ML

Approach diagram

flowchart LR
  UMI["UMI demos<br/>10k+ hours, hand-held gripper"] --> RDT2["RDT2 โ€” Qwen2.5-VL-7B-Instruct<br/>(2 wrist fisheye imgs + language)"]
  RDT2 --> S1["Stage 1: RDT2-VQ<br/>RVQ tokens + cross-entropy"]
  RDT2 --> S2["Stage 2: RDT2-FM<br/>400M RDT flow-matching expert"]
  S2 --> S3["Stage 3: RDT2-UltraFast<br/>one-step distillation"]
  S1 --> ACT["action chunk"]
  S3 --> ACT
  ROBOT_NEW["unseen robot at deploy"] --> RDT2
Loading

Problem

RDT-1B (1.2 B) showed continuous diffusion VLAs can scale across embodiments via a unified action space. The question RDT2 attacks: how far does this scale, and what is the data ceiling? Real-robot teleop data is bottlenecked at hundreds of hours; UMI (universal manipulation interface, hand-held gripper) data scales to 10 000+ hours.

Method

  • Build on a Qwen2.5-VL-7B-Instruct VLM backbone, taking two wrist-view fisheye images + a language instruction as input.
  • Train on 10 000+ hours of UMI demonstrations (one of the largest open-source robot datasets) collected with an enhanced, embodiment-agnostic UMI.
  • A three-stage recipe that aligns the VLM's discrete linguistic knowledge with continuous control:
    • Stage 1 โ€” RDT2-VQ: discretize action chunks with RVQ (Residual Vector Quantization) and train the VLM with cross-entropy. A 0.8 s chunk @ 30 Hz compresses to 27 tokens (โ‰ˆ1/3 of FAST, 1/8 of binning); requires 27 autoregressive passes per chunk. This discrete head also enables RL.
    • Stage 2 โ€” RDT2-FM: attach a 400 M RDT flow-matching action expert (improved RDT model) for continuous action generation without quantization error โ€” one Qwen pass + ~5 expert passes. RDT2-FM-Post adds real-robot data (UR, Franka).
    • Stage 3 โ€” RDT2-UltraFast: distill RDT2-FM into a one-step diffusion policy โ€” one Qwen pass + one expert pass for real-time inference.
  • Evaluate zero-shot transfer to robots not seen in training.

Results

Among the first models to simultaneously zero-shot generalize across the "4U" axes โ€” unseen embodiment, scene, object, and language โ€” since UMI demonstrations are robot-agnostic in their hand-held form. Reported to outperform SOTA baselines such as ฯ€0-FAST and ฯ€0.5 on dexterous, long-horizon, and dynamic downstream tasks. Demos include table tennis (predictive ball-trajectory planning under a 1 m/s arm limit), archery (~100 ms reaction), and fabric folding generalizing to unseen textures/sizes. The public materials emphasize qualitative demos; per-task quantitative tables were not extractable from the project page / arXiv HTML at audit time.

Significance

This is the cleanest scaling-law evidence for the UMI-data โ†’ general policy pathway. UMI's bet is that hand-held demonstrations are an underutilized data source 10โ€“100ร— larger than teleop; RDT2 validates that this is enough to learn a generalist policy with zero-shot transfer. Most relevant predecessors: DexUMI (UMI for dexterous hands), X-VLA (soft-prompt cross-embodiment).

Links

Related pages

โ† Back to CVPR-2026

โš ๏ธ **GitHub.com Fallback** โš ๏ธ