CVPR 2026 DM0 - Heungwoo/research GitHub Wiki
Venue: CVPR 2026 Category: VLA Architecture (Embodied-native pretraining) Trend tag: Trend 1 Team: Dexmal & StepFun (project leads Erjin Zhou, Tiancai Wang) ยท arXiv Feb 16 2026 Backbone: Qwen3-1.7B LLM + perception encoder; ~2B total params with a flow-matching action expert
flowchart TB
subgraph Pre["Stage 1: Pretraining (1.2T tokens)"]
WEB["web text + image"]
DRIVE["driving data"]
EMBO["embodied logs"]
end
Pre --> BB["DM0 VLM backbone<br/>(Qwen3-1.7B + perception encoder)"]
BB --> MID["Stage 2: Mid-Training<br/>add flow-matching action expert<br/>(text + discrete + continuous actions)"]
MID --> POST["Stage 3: Post-Training<br/>specialize per embodiment"]
POST --> ESS["Embodied Spatial Scaffolding<br/>subtask โ goal bbox โ EE traj โ action tokens"]
ESS --> ACT["action prediction"]
Most VLAs are language-pretrained then fine-tuned on robot data. This means the backbone learned its world prior from text and image data that has no embodied structure โ no notion of contact, force, kinematics, or spatial scaffolding. The resulting models are good at language, weak at embodied reasoning.
DM0 uses a three-stage pipeline on a Qwen3-1.7B-based VLM (with a perception encoder), ~2B params total:
- Pretraining (~1.2T tokens): unified large-scale training on heterogeneous data โ (a) web text + image, (b) driving data (autonomous-driving trajectories with rich spatial structure), and (c) embodied logs โ from the start, not as a fine-tune, so semantic knowledge and physical priors are acquired concurrently.
- Mid-Training: attaches a flow-matching action expert atop the VLM and jointly supervises text tokens, discrete action tokens, and continuous actions. A hybrid gradient strategy decouples the action-expert gradients from the VLM on embodied data (they are not backpropagated into the backbone) so the VLM keeps learning on non-embodied data and does not erode its general knowledge.
- Post-Training: specializes the model per target embodiment while retaining dialogue ability.
Embodied Spatial Scaffolding is a hierarchical supervision / structured information bottleneck applied in mid/post-training (not a pretraining objective). The model sequentially predicts subtask decomposition โ goal bounding boxes โ end-effector trajectory โ discrete action tokens, progressively constraining the action hypothesis space. DM0 unifies manipulation and navigation (navigation trajectories from Habitat are included).
State-of-the-art on the RoboChallenge Table30 benchmark with only ~2B params:
| Setting | Model | Success rate | Task score |
|---|---|---|---|
| Specialist | DM0 | 62.00% | โ |
| Specialist | GigaBrain-0.1 | 51.67% | โ |
| Specialist | Spirit-v1.5 | 51.00% | โ |
| Specialist | ฯ0.5 | 42.67% | โ |
| Generalist | DM0 | 37.3% | 49.08 |
| Generalist | ฯ0.5 | 17.67% | 31.27 |
| Generalist | ฯ0 | 9.0% | 20.22 |
DM0 is the most explicit "embodiment-from-day-one" pretraining recipe published to date. Sits philosophically opposite to the standard "freeze a web VLM, bolt on an action head" pattern (GR00T series, ฯ0.7). If the trend holds, expect a 2026โ2027 split between web-pretrained โ robot-fine-tuned (current dominant paradigm) and jointly embodiment-pretrained (DM0's bet).
- arXiv: 2602.14974
- Code: Dexmal/dexbotic
โ Back to CVPR-2026