CoRL 2026 VLA Feedback - Heungwoo/research GitHub Wiki

CoRL 2026 โ€” VLA-Feedback: Real-Time Feedback Denoising for Responsive VLAs

Venue: CoRL 2026 (Austin, TX, Nov 9โ€“12). Paper: arXiv 2609.21022 ("Catch Me If You Can: Real-Time Feedback Denoising for Responsive VLAs"). Representative of: reactive execution inside the action chunk โ€” keep the plan, correct each action on the final denoising step with fresh vision. Companions: Real-Time Execution ยท CoRL 2026 survey.

Overview of VLA-Feedback: a traditional VLA executes a generated action chunk open-loop, while VLA-Feedback keeps the VLM-DiT planner in a slow pathway and corrects each action against the latest observation via a lightweight Feedback Denoising Module (figure from the authors, arXiv 2609.21022, ยฉ the authors)

1. Problem

Diffusion-based VLAs (e.g. GR00T-style VLM-DiT planners) infer an action chunk and then execute it open-loop until the next inference. During that window the model is blind: if an object moves, contact dynamics shift, or the scene evolves mid-chunk, the pre-generated actions fire against stale observations and miss the (now-moved) target. Simply re-running the heavy vision-language planner every step is too slow to close the loop.

2. Method

VLA-Feedback is a two-timescale design. The slow pathway is the unchanged VLM-DiT diffusion planner, run at low frequency to produce action chunks. The insight is to retain the final denoising step as a lightweight feedback interface rather than finishing denoising before acting: a small Feedback Denoising Module โ€” a lightweight transformer that fuses cached action representations with the current visual observation โ€” predicts an observation-conditioned denoising velocity update and applies one denoising update (with a step size) to refine each near-final action just before it fires. This gives high-frequency intra-chunk correction without rerunning the planner, preserving the expressiveness of the full diffusion plan.

3. Results

  • Static LIBERO: matches GR00T โ€” LIBERO-Goal 95.5% vs 97.5%, LIBERO-Object 92% vs 92% (no regression on static tasks).
  • Dynamic simulation: 27.5% โ†’ 85.0% average success vs the GR00T baseline.
  • Real robot: 51% โ†’ 73% average success.

4. Why it matters

Reactivity is usually framed as an inference-speed problem (chunk faster, or interpolate between chunks). VLA-Feedback reframes it: the last denoising step is already a natural correction interface, so a cheap module can steer each action with fresh vision while the expensive planner keeps running at its own slow rate. It preserves static-task quality while dramatically improving performance on moving targets โ€” a practical recipe for closing the loop inside the chunk.

Limitations (reviewer): It is a local refinement โ€” it depends on a reasonable initial plan; when chunks start far from a valid solution, or the target moves near contact, one-step feedback denoising may lack the workspace to recover. Real-world gains are also bounded by occlusion, actuation delay, and observation noise.

5. Links

โ† Back to CoRL 2026 survey ยท Home