CoRL 2026 StressDream - Heungwoo/research GitHub Wiki
CoRL 2026 ā StressDream: Steering Video World Models for Robust Policy Evaluation
Venue: CoRL 2026 (Austin, TX, Nov 9ā12) Ā· CMU IntentLab. Paper: arXiv 2606.00267. Representative of: world-model-as-stress-tester ā don't imagine the likely future, imagine plausible failures to expose where the policy breaks. Companions: World Models Ā· VLA Evaluation Ā· CoRL 2026 survey.

1. Problem
Video world models (WMs) are increasingly used to evaluate policies by rolling out imagined futures. But nominal sampling reproduces the likely future, so rare-but-catastrophic failure modes (collisions, task breakdowns) almost never surface without a prohibitive number of samples. StressDream asks the WM to instead conjure high-impact yet still plausible outcomes on demand, so evaluation finds where a policy breaks and training can then patch it.
2. Method
The initial Gaussian noise of a diffusion video WM is treated as a control variable and optimized by gradient ascent, ε ā ε + Ī·ā_ε C(o), steering the imagined rollout toward a target outcome. Two objectives shape the criterion C:
- Semantic guidance (VLM): an inference-time text prompt l names the target failure (e.g. "the car collides"). A VLM scores the video, C_sem = log p(yes|o,l) ā log p(no|o,l). Gradients use a score-distillation shortcut (ā_ε C ā βā_o C), so only the differentiable criterion is backpropagated, not the full denoiser.
- Plausibility constraint: keeps the optimized noise inside the Gaussian typical set via norm concentration (āεāā ā āD), blockwise isotropy, and spectral-whiteness terms ā so the imagined failures stay realistic rather than adversarial garbage.
Outer loop (policy improvement): trajectories where steering finds a failure are down-weighted (0.1 vs 1.0) in a weighted-regression fine-tune of the policy, teaching it to propose actions that avoid the discovered failure modes.
3. Results
- Driving (Vista WM): predicts 25 future 576Ć1024 frames from waypoint actions (noise dim Dā922k). Evaluated on 100 safety-critical PAI-AV pairs across 8 event categories plus 200 imminent-collision cases. Beats Best-of-N sampling on both target-alignment and video-quality; higher recall for surfacing failure events than nominal generation.
- Manipulation (Ctrl-World on DROID): 5 future frames, 3 views at 192Ć320, joint-position actions (Dā58k), 6 contact-rich tasks. Robustly fine-tuning Ļā.ā on discovered failures lifts task success from 39% ā 71%.
- Dubins-car case study: with a 20% stochastic control-flip; StressDream keeps a high true-positive and true-negative rate, whereas dropping the plausibility objective collapses the true-negative rate (implausible "failures" imagined).
4. Why it matters
Turns a generative WM from a passive rollout engine into an active stress-tester: text-specifiable, gradient-steered failure discovery that plugs straight into policy fine-tuning ā a cheap alternative to hand-scripted adversarial scenarios or massive random sampling.
Limitations (reviewer): semantic signal is only as good as the VLM's ability to recognize the named failure; the plausibility constraint is a proxy for in-distribution-ness, not a guarantee; steering targets must be enumerated as text prompts, so unknown failure modes stay unimagined.
5. Links
- arXiv 2606.00267
- Survey: CoRL 2026 Ā· Related: World Models Ā· VLA Evaluation
ā Back to CoRL 2026 survey Ā· Home