ICML 2026 From Noise to Intent - Heungwoo/research GitHub Wiki

From Noise to Intent — ResVLA anchors generative policies on predicted intent instead of starting from pure noise

Venue: ICML 2026 (Poster) Category: Diffusion-Flow Policy / VLA Architecture Traction (2026-06): 0 citations (arXiv)

Paradigm comparison: Generation-from-Noise vs. Refinement-from-Intent (Figure 1 from Zhong et al., 2026)

Problem

Generative VLA policies typically follow a "Generation-from-Noise" paradigm: actions are sampled by transporting uninformative Gaussian noise to the expert action manifold. This ignores the fundamental spatiotemporal scale mismatch between high-level cognition and low-level control, causing representation inefficiency, weak conditioning, and a pathology the authors call Loss Collapse — when the generative source is condition-independent, trajectories fail to align with semantic instructions during optimization. The result is policies that overfit to visual-textual correlations on benchmarks like LIBERO yet collapse under out-of-distribution language, embodiment, or layout perturbations.

Method

ResVLA shifts the paradigm to "Refinement-from-Intent." It decomposes robotic motion via a Discrete Cosine Transform (DCT) into two complementary subspaces: a Semantic Subspace spanned by the lowest k frequency modes (a learnable cutoff) that captures the smooth, low-frequency global trajectory — the deterministic anchor — and an orthogonal Execution Subspace of high-frequency jitter that captures local contact dynamics — the stochastic residual. An Intent Anchoring Module uses VLM features to directly regress the low-frequency component, constructing a condition-dependent source distribution pā‚€(x|c) that satisfies the theory's requirement for a condition-dependent origin (avoiding Loss Collapse). A flow-matching expert then learns a Residual Diffusion Bridge: a short transport path from this anchor to the full ground-truth action, focusing strictly on refining high-frequency execution details. Because transport starts from a task-aligned anchor rather than noise, the path has minimal curvature ("Path Straightening"), enabling few-step inference.

Overview of the two-stage ResVLA framework: Intent Anchoring then Residual Bridging (Figure 2 from Zhong et al., 2026)

Results

On standard LIBERO, ResVLA is competitive with SOTA. On the harder LIBERO-Plus robustness suite it is far more resilient: 88.5% success on diverse language instructions (↑7.5% over the best baseline) where OpenVLA plummets to 23.0%; 59.9% in the Robot (embodiment) setting where π₀ collapses to 6.0%; and a SOTA 79.0% in the Layout setting. Camera-viewpoint perturbation remains the hardest axis. Efficiency: trained from scratch (30k steps, no pre-training), ResVLA converges faster than π₀ and stays robust at high dropout (p=0.2) where π₀ saturates; it reaches ~85% at NFE=4 and still hits 70% at a single function evaluation (NFE=1). Cross-embodiment (SimplerEnv): 68.4% avg on Google Robot (vs. π₀ 58.8%, OpenVLA 34.3%, RT-1-X 42.4%) and 57.9% on WidowX, beating Octo-Base, OpenVLA, CogACT and SpatialVLA. Real-world: competitive on a contact-rich three-stage ALOHA bimanual task (Pick Cup → Handover → Placement) where errors accumulate across stages.

Significance

ResVLA challenges the "noise-to-action" orthodoxy by grounding generation in a predicted semantic anchor, directly attacking the Loss Collapse pathology with frequency-domain structure. The payoff is strong robustness to language/embodiment/layout shift and near-single-step inference, suggesting spectral anchoring is a principled route to efficient, generalizable generative control.

Links

← Back to ICML-2026