CoRL 2026 StressDream - Heungwoo/research GitHub Wiki

CoRL 2026 — StressDream: Steering Video World Models for Robust Policy Evaluation

Venue: CoRL 2026 (Austin, TX, Nov 9–12) Ā· CMU IntentLab. Paper: arXiv 2606.00267. Representative of: world-model-as-stress-tester — don't imagine the likely future, imagine plausible failures to expose where the policy breaks. Companions: World Models Ā· VLA Evaluation Ā· CoRL 2026 survey.

StressDream steers a diffusion video world model's initial noise toward a text-specified high-impact outcome, then trains on the discovered failures (figure from Seo et al., arXiv 2606.00267, Ā© the authors)

1. Problem

Video world models (WMs) are increasingly used to evaluate policies by rolling out imagined futures. But nominal sampling reproduces the likely future, so rare-but-catastrophic failure modes (collisions, task breakdowns) almost never surface without a prohibitive number of samples. StressDream asks the WM to instead conjure high-impact yet still plausible outcomes on demand, so evaluation finds where a policy breaks and training can then patch it.

2. Method

The initial Gaussian noise of a diffusion video WM is treated as a control variable and optimized by gradient ascent, ε ← ε + Ī·āˆ‡_ε C(o), steering the imagined rollout toward a target outcome. Two objectives shape the criterion C:

  • Semantic guidance (VLM): an inference-time text prompt l names the target failure (e.g. "the car collides"). A VLM scores the video, C_sem = log p(yes|o,l) āˆ’ log p(no|o,l). Gradients use a score-distillation shortcut (āˆ‡_ε C ā‰ˆ Ī²āˆ‡_o C), so only the differentiable criterion is backpropagated, not the full denoiser.
  • Plausibility constraint: keeps the optimized noise inside the Gaussian typical set via norm concentration (‖ε‖₂ ā‰ˆ √D), blockwise isotropy, and spectral-whiteness terms — so the imagined failures stay realistic rather than adversarial garbage.

Outer loop (policy improvement): trajectories where steering finds a failure are down-weighted (0.1 vs 1.0) in a weighted-regression fine-tune of the policy, teaching it to propose actions that avoid the discovered failure modes.

3. Results

  • Driving (Vista WM): predicts 25 future 576Ɨ1024 frames from waypoint actions (noise dim Dā‰ˆ922k). Evaluated on 100 safety-critical PAI-AV pairs across 8 event categories plus 200 imminent-collision cases. Beats Best-of-N sampling on both target-alignment and video-quality; higher recall for surfacing failure events than nominal generation.
  • Manipulation (Ctrl-World on DROID): 5 future frames, 3 views at 192Ɨ320, joint-position actions (Dā‰ˆ58k), 6 contact-rich tasks. Robustly fine-tuning π₀.ā‚… on discovered failures lifts task success from 39% → 71%.
  • Dubins-car case study: with a 20% stochastic control-flip; StressDream keeps a high true-positive and true-negative rate, whereas dropping the plausibility objective collapses the true-negative rate (implausible "failures" imagined).

4. Why it matters

Turns a generative WM from a passive rollout engine into an active stress-tester: text-specifiable, gradient-steered failure discovery that plugs straight into policy fine-tuning — a cheap alternative to hand-scripted adversarial scenarios or massive random sampling.

Limitations (reviewer): semantic signal is only as good as the VLM's ability to recognize the named failure; the plausibility constraint is a proxy for in-distribution-ness, not a guarantee; steering targets must be enumerated as text prompts, so unknown failure modes stay unimagined.

5. Links

← Back to CoRL 2026 survey Ā· Home