Review Cortex2 - Heungwoo/research GitHub Wiki
Model: Cortex 2.0 — a VLA augmented with a world-model foresight-planning layer (modular WAM+VLA hybrid) · Sereact GmbH (Stuttgart) Paper: "Cortex 2.0: Grounding World Models in Real-World Industrial Deployment" — arXiv 2604.20246 (Apr 22 2026) · blog Status: arXiv preprint — results are company/deployment-reported from Sereact's production fleet, not independently replicated. Filed under Latest Papers. Backing: $110M Series B (2026), scaling Cortex 2 + US expansion.
The blog's headline: "Today's Cortex sees and picks; Cortex 2.0 thinks first, then acts." It is the modular/bolted-on exemplar of the WAM+VLA hybrid (NVIDIA WAM thesis §2.6) — a world model added as a foresight-planning layer in front of a VLA action head. Companion: VLA Architectures §4.2b · World Models · MOTUS · Being-H0.7.
- A world model as a planning layer bolted onto a VLA. Four hierarchical stages: (1) High-level VLM (a 2B-VLM) encodes the scene into task context → (2) World Model generates k candidate future trajectories in visual latent space via flow matching → (3) PRO (Process-Reward Operator) scores each → (4) flow-matching action head commits to the best branch and replans in real time at 30 Hz.
-
Scoring by progress, risk, termination. PRO uses frozen heads (trained on deployment data): Progress (Δ value V_φ), Risk (failure probability, penalizing high-speed contact / compression / edge impacts), Termination (success likelihood). Composite
S_j = Δp − λ·ρ + β·d; pickargmax, then feed a binarized advantage signal into the policy. - Adaptive foresight = a compute dial. More rollouts when failure is expensive (packing, fragile placement), fewer when recovery is cheap (regrasp). Success rises 0.962 (k=1) → 0.996 (k=30) while step latency grows 310 ms → 9,200 ms.
- Strong deployment-reported wins over π0.5 / Diffusion Policy / RDT-2 across four industrial tasks — with 0 human interventions where baselines needed dozens.
- It makes "world-model foresight" a deployable planning layer, not a research toy. Cortex 2.0 keeps a fast VLA action head and only scores imagined branches — a pragmatic answer to the WAM latency problem (Review-WAM-vs-VLA-Robustness): imagination is used for selection, and the compute is a tunable dial per task.
- It is the clearest industrial datapoint for the modular hybrid. Contrast the unified hybrids (MOTUS, Being-H0.7, Cosmos 3): Cortex 2.0 bolts the world model on as a process-reward planner, closer in spirit to test-time search (VLA-Reasoner) than to a single MoT.
- Trained on real deployment at scale (>10M episodes / >25k h of warehouse operation) — the data regime academic hybrids can't access.
flowchart LR
O[observation] --> VLM[High-level VLM · 2B<br/>task context s_t]
VLM --> WM[World model<br/>k candidate futures in visual latent<br/>flow matching, ODE σ:0→1]
WM --> PRO[PRO scorer<br/>progress − λ·risk + β·termination]
PRO -->|argmax → binarized advantage I_t| AH[Flow-matching action head]
AH ==>|30 Hz, real-time replan| ACT[action]
-
Candidate generation. Conditioned on current latent
z_tand task contexts_t, each of k rollouts starts from a distinct noise drawξ⁽ʲ⁾∼𝒩(0,I), integrated by ODE from σ=0→1 over horizonH_wm. -
PRO (Process-Reward Operator). Three frozen heads score each rollout: progress
Δp, riskρ, terminationd→S_j = Δp − λρ + βd. The winning branch's advantage is binarizedI_t∈{0,1}and fed to the policy as conditioning. -
Embodiment adaptation is handled entirely by the action heads (learned projections
W_z, W_I), so the planner transfers across platforms. - Data: deployment >10M episodes / >25k h (warehouse), teleop ~40k/~400 h, open-source ~970k/~2k h (OXE, BridgeData V2, DROID), synthetic RoboCasa ~20k. Baselines trained at equal 200 GPU-hour budgets.
| Task (dual-arm unless noted) | Cortex 2.0 | π0.5 | Diffusion Policy | RDT-2 |
|---|---|---|---|---|
| Single-arm pick-and-place (16 trials) | 0.98 · 20 s · 0 | 0.7 · 49 s · 2 | 0.56 · 53 s · 4 | 0.4 · 63 s · 7 |
| Sorting items & trash (8.7k ep) | 0.95 · 700 s · 0 | 0.61 · DNF · 53 | 0.47 · DNF · 59 | 0.18 · DNF · 95 |
| Sorting screws (3.1k ep) | 0.98 · 180 s · 0 | 0.4 · DNF · 24 | 0.2 · DNF · 16 | 0.0 · DNF · 50 |
| Shoebox unpacking (2.9k ep) | 0.96 · 58 s · 0 | 0.6 · 103 s · 5 | 0.12 · 52 s · 9 | 0.0 · 62 s · 10 |
(DNF = did not complete; per-operation success for the multi-item tasks.)
Significance. Cortex 2.0 shows world-model foresight-as-planning working on a real industrial fleet with zero interventions — the strongest deployment evidence yet for the modular WAM+VLA hybrid, and a clean separation of imagination (selection) from control (fast action head).
Limitations.
- Company/deployment-reported, not peer-reviewed or independently replicated. Results come from Sereact's own production stack and tasks.
- Proprietary data & environments. The >10M-episode deployment corpus is closed; generalization outside Sereact deployments is untested.
- Latency–foresight tradeoff is steep. k=30 reaches 0.996 but at 9.2 s/step — high-k foresight is viable only where failure cost dominates cycle time.
-
Per-task tuning. The advantage-binarization threshold
ϵ(s_t)and the k/H_wm budget are hand-set per task. - No standardized benchmark. Baselines are re-trained in-house at matched compute, but there's no shared LIBERO/RoboTwin comparison.
- Paper: arXiv 2604.20246 · blog: sereact.ai/posts/cortex-2
- VLA Architectures §4.2b (WAM×VLA hybrids) · World Models · WAM vs VLA Robustness
- Sibling hybrids: MOTUS · Being-H0.7 · DYNA-2 · Cosmos 3 / NVIDIA WAM
- Latest Papers
← Back to Latest Papers · Home · Reviews