CVPR 2026 DM0 - Heungwoo/research GitHub Wiki

DM0 โ€” An Embodied-Native VLA towards Physical AI

Venue: CVPR 2026 Category: VLA Architecture (Embodied-native pretraining) Trend tag: Trend 1 Team: Dexmal & StepFun (project leads Erjin Zhou, Tiancai Wang) ยท arXiv Feb 16 2026 Backbone: Qwen3-1.7B LLM + perception encoder; ~2B total params with a flow-matching action expert

Approach diagram

flowchart TB
  subgraph Pre["Stage 1: Pretraining (1.2T tokens)"]
    WEB["web text + image"]
    DRIVE["driving data"]
    EMBO["embodied logs"]
  end
  Pre --> BB["DM0 VLM backbone<br/>(Qwen3-1.7B + perception encoder)"]
  BB --> MID["Stage 2: Mid-Training<br/>add flow-matching action expert<br/>(text + discrete + continuous actions)"]
  MID --> POST["Stage 3: Post-Training<br/>specialize per embodiment"]
  POST --> ESS["Embodied Spatial Scaffolding<br/>subtask โ†’ goal bbox โ†’ EE traj โ†’ action tokens"]
  ESS --> ACT["action prediction"]
Loading

Problem

Most VLAs are language-pretrained then fine-tuned on robot data. This means the backbone learned its world prior from text and image data that has no embodied structure โ€” no notion of contact, force, kinematics, or spatial scaffolding. The resulting models are good at language, weak at embodied reasoning.

Method

DM0 uses a three-stage pipeline on a Qwen3-1.7B-based VLM (with a perception encoder), ~2B params total:

  1. Pretraining (~1.2T tokens): unified large-scale training on heterogeneous data โ€” (a) web text + image, (b) driving data (autonomous-driving trajectories with rich spatial structure), and (c) embodied logs โ€” from the start, not as a fine-tune, so semantic knowledge and physical priors are acquired concurrently.
  2. Mid-Training: attaches a flow-matching action expert atop the VLM and jointly supervises text tokens, discrete action tokens, and continuous actions. A hybrid gradient strategy decouples the action-expert gradients from the VLM on embodied data (they are not backpropagated into the backbone) so the VLM keeps learning on non-embodied data and does not erode its general knowledge.
  3. Post-Training: specializes the model per target embodiment while retaining dialogue ability.

Embodied Spatial Scaffolding is a hierarchical supervision / structured information bottleneck applied in mid/post-training (not a pretraining objective). The model sequentially predicts subtask decomposition โ†’ goal bounding boxes โ†’ end-effector trajectory โ†’ discrete action tokens, progressively constraining the action hypothesis space. DM0 unifies manipulation and navigation (navigation trajectories from Habitat are included).

Results

State-of-the-art on the RoboChallenge Table30 benchmark with only ~2B params:

Setting Model Success rate Task score
Specialist DM0 62.00% โ€”
Specialist GigaBrain-0.1 51.67% โ€”
Specialist Spirit-v1.5 51.00% โ€”
Specialist ฯ€0.5 42.67% โ€”
Generalist DM0 37.3% 49.08
Generalist ฯ€0.5 17.67% 31.27
Generalist ฯ€0 9.0% 20.22

Significance

DM0 is the most explicit "embodiment-from-day-one" pretraining recipe published to date. Sits philosophically opposite to the standard "freeze a web VLM, bolt on an action head" pattern (GR00T series, ฯ€0.7). If the trend holds, expect a 2026โ€“2027 split between web-pretrained โ†’ robot-fine-tuned (current dominant paradigm) and jointly embodiment-pretrained (DM0's bet).

Links

Related pages

โ† Back to CVPR-2026

โš ๏ธ **GitHub.com Fallback** โš ๏ธ