Review Dyna2 - Heungwoo/research GitHub Wiki

In-Depth Review — DYNA-2: A World-Action Model with a Human-to-Robot Scaling Law

Model: DYNA-2 — World-Action Model (WAM) · Dyna Robotics Announced: August 10, 2026 (company launch). Primary source: Dyna Robotics' own technical writeup — Dyna-2: A 1-Million-Hour Scaling Law for World-Action Models (architecture, power-law fits with R², ablations, per-task results). Self-published; not peer-reviewed; no arXiv preprint or released weights/code. Founders: Lindon Gao, York Yang (prev. Caper AI, $350M exit), Jason Ma (ex-DeepMind) · Backers: CRV, First Round Capital Press coverage: PR Newswire · TechTimes Status: industry release — filed under Latest Papers.

⚠️ Sourcing caveat. The numbers below come from Dyna Robotics' own technical report (dyna.co/dyna-2) — which, unlike a bare press release, does disclose the architecture, published power-law fits with R² values, ablations, and per-task results. But it is self-published and not peer-reviewed: no independent replication, no released weights/code, no shared benchmark, and model size / training compute are not disclosed. Treat the figures as credible-but-unverified vendor results.

Companion reviews: World Models · DreamZero · ω-0 · Human Video → Robot Transfer · EgoScale · WAM vs VLA Robustness.


1. TL;DR

  1. A World-Action Model trained entirely on human egocentric video — no robot data in pre-training, using a mixture-of-transformers trained with flow matching to jointly predict the next video frame and next action. Pre-training: >1,000,000 h of head-mounted egocentric human video (cooking, tidying, folding, assembling), with 3D hand-pose tracks as pseudo-action supervision. Model size undisclosed.
  2. The one structural fact that reframes everything: video prediction is a co-training objective that is dropped at inference. The action head never takes the video latent as input, and at inference "the policy neither generates nor attends to predicted future video" — DYNA-2 is a reactive policy that keeps the world-model representation benefit without paying video-generation latency (§3).
  3. A demonstrated human-to-robot scaling law with published fits. Across four nested scales (1k→10k→100k→1M h): held-out human MSE = 0.0691·D⁻⁰·⁰¹⁸⁴ (R²=0.919); and the substantive claim — zero-shot robot-transfer action MSE = 0.306·D⁻⁰·⁰⁷¹³ (R²=0.884), 0.195→0.117 with no robot data in pre-training. ~50× EgoScale's 20,854 h.
  4. On-robot after fine-tuning (14 tasks): 20→28→45→53% across 1k→1M h; precision skills emerge at scale (lockbox key-turning 0% up to 100k h → 90% at 1M h). vs its prior VLA DYNA-1: 87% vs 46% zero-shot at a customer site, 1.55× success, language-following 35→67→96% via video co-training. 13 min of teleop data sufficed to fine-tune bottle-cap opening.

2. Why this matters

  • The most aggressive bet yet on the WAM + human-video thesis — with a published curve. DYNA-2 extends the EgoScale scaling axis ~50× on pure human video, no robot data in pre-training, and publishes the power-law fits (R² ≈ 0.88–0.93). It is the strongest form yet of "abundant human video substitutes for scarce robot data" (human-video fork).
  • It quietly resolves the WAM latency tension. The field's standing complaint is that world-action models are too slow for the control loop (WAM inference measured at 4.8×–83× π0.5 in the robustness study). DYNA-2's decoupled design — world-model as a co-training signal, dropped at inference — is a concrete answer (§3, §5).
  • The scaling law is documented, not just asserted. The remaining gap versus EgoScale is provenance (peer review, released weights, independent replication), not disclosure.

3. Model architecture in detail

DYNA-2 is a mixture-of-transformers: every modality (video, action, proprioception, text) is tokenized separately, gets its own stack of DiT layers, and the stacks exchange information through attention. The asymmetry between the streams is the design.

flowchart LR
  subgraph IN[Inputs]
    V[Observed video context]
    T[Text instruction]
    P[Proprioception]
  end
  V -->|causal mask| VID[Video DiT stack<br/>deep]
  T -->|cross-attn| VID
  P --> ACT[Action DiT stack<br/>shallow, bidirectional]
  VID -.->|early layers only,<br/>context tokens| ACT
  VID --> VH["Video flow-matching head<br/>(predict future frames)"]
  ACT --> AH["Action flow-matching head<br/>(predict action chunk)"]
  VH -. dropped at inference .-> X((✗))
  AH ==>|inference output| OUT[Action chunk]
  classDef drop fill:#eee,stroke:#999,color:#999;
  class X,VH drop
Loading

Streams.

  • Video stream — a video-diffusion backbone in latent space z_t. Video tokens use causal masking (past→future) and cross-attend to text, so language conditions the world prediction.
  • Action stream — supervised by pseudo-actions derived from 3D hand-pose tracks: wrist poses → end-effector trajectories, plus a continuous grasp signal from the thumb–index aperture. This is how label-free human video yields action supervision. Action tokens use bidirectional self-attention and attend to the observed (context) video tokens; the action transformer is deliberately shallow and joins the video stream only at early layers.
  • Proprioception is tokenized and fed straight to the action transformer.
  • Text → video only — language cross-attends to video tokens but does not directly influence action tokens; instruction-following is mediated through the world representation.

Training — joint flow matching. Both streams are trained by flow matching (z_t = t·z + (1−t)·ε_z, a_t = t·a + (1−t)·ε_a) with a co-training loss L_co = E‖u_θ^vid − (z−ε_z)‖² + λ·E‖u_θ^act − (a−ε_a)‖². The pivotal choice: the action head never takes z_t as an argument. Video and action objectives share the early representation but are decoupled at the output — the model does not condition actions on generated future frames.

Inference — reactive, video dropped. Per the report, "the policy neither generates nor attends to predicted future video at inference time." It consumes current observations + proprioception and emits an action chunk directly. Video prediction is a representation-shaping co-training objective, not an inference-time imagination loop. This is exactly what lets a "world-action model" stay real-time reactive.

One-step video distillation (for the generative use of the same backbone). A continuous-curriculum distillation (coupled fast student-update / slow target-advance) compresses the 100-step teacher into a one-step student: ~90× faster (10,203→110 ms on H100), FVD 121 vs teacher 80, 75% motion retention, flicker 1.94 (real 2.37). Authors note it "does not yet match the full-precision teacher."


4. What is (and isn't) disclosed

Aspect Company statement Disclosed?
Architecture Mixture-of-transformers; per-modality DiT; flow matching; video causal-masked + text cross-attn, action bidirectional + shallow, proprio→action; action head never sees z_t; reactive inference ✅ design disclosed
Model size / compute — ❌ not disclosed
Pre-training data >1,000,000 h egocentric human video; hand-pose (wrist + thumb–index) pseudo-actions; no robot data ✅
Scaling study Nested 1k/10k/100k/1M h; power-law fits w/ R² for human MSE/acc and zero-shot robot MSE/acc ✅ curves + fits
Zero-shot human→robot transfer action MSE 0.195→0.117 (R²=0.884); [email protected] 0.067→0.159 ✅
On-robot (14 tasks, fine-tuned) mean 20→28→45→53%; lockbox 0→90%, bottle-cap 10→50%, drink 58→83% ✅
vs DYNA-1 87% vs 46%; 1.55× success, 1.12× grade; language 35→67→96% ✅ (self-described lower bound)
Peer review · replication · weights · code · shared benchmark — ❌ none

5. How DYNA-2 differs from other WAMs

The most useful axis is where the world model lives at inference — does the policy actually generate/consume future frames to act (slow), or is future-prediction only a training-time signal (fast)?

System Predicts Future-pred in the action path at inference? Pre-training data Scaling evidence Openness / cost
DYNA-2 (this) pixel video (co-train) + action No — reactive, video dropped >1M h human, no robot published fits, R²≈0.88–0.93 closed; real-time
DreamZero 14B pixel video + action co-generated Yes (frames in the action path) ~500 h robot + video backbone-size ablation (14B>5B) open; ~7 Hz
ω-0 latent future (reconstruction-free) No — video dropped from action path human motion + 40 h ω-HOME task-suite results preprint; real-time
Cosmos Policy latent frame (action+future+value) Yes (co-generated) robot 185-traj efficiency ~390 ms (6.2× π0.5)
DreamGen / mimic-video video → IDM decode offline (data factory) robot + video log-linear synthetic scaling offline
DreamDojo continuous-latent-action WM (planner/latent) 44k h egocentric human — real-time 10.9 FPS
EgoScale ego-video pretraining (not a WAM) n/a (representation) 20,854 h ego peer-reviewed, R²=0.9983 paper

Three concrete differences:

  1. Scale + purity of the data fork. At >1M h, human-only, DYNA-2 is the extreme of the human-video fork. Most WAMs co-train robot+video; DYNA-2 learns actions from human video alone via hand-pose pseudo-actions (wrist → EE, thumb–index → grasp).
  2. The world model is dropped at inference (the reactive corner). DYNA-2 and ω-0 both discard future-prediction at action time; DreamZero and Cosmos-Policy keep it (co-generate, and pay 390 ms–seconds). DYNA-2 differs from ω-0 in substrate: DYNA-2 trains pixel video prediction as the co-training signal (then drops it); ω-0 predicts latent futures. In World-Models taxonomy terms, DYNA-2 uses a Family-A substrate (pixel video) in an E1 role (auxiliary representation) — a hybrid the taxonomy didn't cleanly name.
  3. The deliverable is a fitted transfer law, not an ablation. DreamZero's evidence is a backbone-size ablation; Cosmos-Policy's is data efficiency. DYNA-2's headline is a human→robot scaling law across four orders of magnitude — closest in spirit to EgoScale (peer-reviewed, 20,854 h) but 50× scale and self-published.

6. Key insights

  1. Future prediction is a scaling enabler, not just an accuracy bump. The joint (video+action) recipe beat action-only on 39 of 39 tasks at every action scale. Action-only "overfits severely and unpredictably as data scales"; the trend reverses only once copious human video is added for video prediction. World-modeling is what makes the data pay off.
  2. Human video dissolves the teleop bottleneck. Teleop "has to be deliberately produced, which bounds pre-training"; egocentric human video exists at ~unbounded scale and already encodes how scenes evolve, how objects respond to contact, and how a hand interacts. Hand-pose pseudo-actions are the bridge that makes label-free video trainable.
  3. World-modeling can be nearly free at inference. Because the action head never sees z_t and video is dropped at inference, DYNA-2 captures the representation benefit without the latency — a direct answer to the "WAMs are too slow for the control loop" tension (Review-World-Models §6.3).
  4. Skills emerge at thresholds, not smoothly. The aggregate curve is smooth, but individual capabilities switch on: lockbox key-turning is 0% up to 100k h and 90% at 1M h; there is an inflection from 10k→100k h. This argues for scale over task-specific engineering — but makes per-task extrapolation risky.
  5. The authors are candid about the caveats. The DYNA-1 comparison is a self-described lower bound ("the entire experimental pipeline is tuned for our VLAs"); compute/model-size scaling is "left for future work" (so data-vs-capacity is unseparated); and "video training dilutes action-learning gradients" — on-embodiment, enough action data alone suffices, so video helps most for transfer and scaling, not in-domain accuracy.

7. Open questions / limitations

  1. Self-published, not peer-reviewed, not replicated. Every number is from Dyna's own eval on Dyna's own 14-task suite; no third party has reproduced the 1M-hour transfer curve.
  2. Model size and compute undisclosed (author-acknowledged). Without params/FLOPs you can't attribute gains to data scale (the claim) vs co-varying model/compute scale — the authors explicitly defer this.
  3. No released weights, code, or shared benchmark. Not reproducible; no comparison to π, GR00T, or Skild on common ground (only to their own VLA, and by their own admission favorably).
  4. Emergence cuts both ways. Threshold skills (0→90%) are impressive but sensitive to task selection and hard to extrapolate; small fitted exponents (human MSE ∝ D⁻⁰·⁰¹⁸⁴) mean most of the headline comes from these jumps.
  5. "No robot data in pre-training" still needs robot data to deploy. Fine-tuning ("hours", "13 minutes") is where action grounding enters; the purity claim is about the prior.
  6. Video-model failure modes likely inherited. The one-step distillation trades fidelity (FVD 121 vs 80, 75% motion retention) for 90× speed; the report has no failure-case section, so contact-rich and long-horizon robustness are untested externally.

8. Links & related pages

← Back to Latest Papers · Home · Reviews

⚠️ **GitHub.com Fallback** ⚠️