CoRL 2026 VLA Feedback - Heungwoo/research GitHub Wiki
CoRL 2026 โ VLA-Feedback: Real-Time Feedback Denoising for Responsive VLAs
Venue: CoRL 2026 (Austin, TX, Nov 9โ12). Paper: arXiv 2609.21022 ("Catch Me If You Can: Real-Time Feedback Denoising for Responsive VLAs"). Representative of: reactive execution inside the action chunk โ keep the plan, correct each action on the final denoising step with fresh vision. Companions: Real-Time Execution ยท CoRL 2026 survey.

1. Problem
Diffusion-based VLAs (e.g. GR00T-style VLM-DiT planners) infer an action chunk and then execute it open-loop until the next inference. During that window the model is blind: if an object moves, contact dynamics shift, or the scene evolves mid-chunk, the pre-generated actions fire against stale observations and miss the (now-moved) target. Simply re-running the heavy vision-language planner every step is too slow to close the loop.
2. Method
VLA-Feedback is a two-timescale design. The slow pathway is the unchanged VLM-DiT diffusion planner, run at low frequency to produce action chunks. The insight is to retain the final denoising step as a lightweight feedback interface rather than finishing denoising before acting: a small Feedback Denoising Module โ a lightweight transformer that fuses cached action representations with the current visual observation โ predicts an observation-conditioned denoising velocity update and applies one denoising update (with a step size) to refine each near-final action just before it fires. This gives high-frequency intra-chunk correction without rerunning the planner, preserving the expressiveness of the full diffusion plan.
3. Results
- Static LIBERO: matches GR00T โ LIBERO-Goal 95.5% vs 97.5%, LIBERO-Object 92% vs 92% (no regression on static tasks).
- Dynamic simulation: 27.5% โ 85.0% average success vs the GR00T baseline.
- Real robot: 51% โ 73% average success.
4. Why it matters
Reactivity is usually framed as an inference-speed problem (chunk faster, or interpolate between chunks). VLA-Feedback reframes it: the last denoising step is already a natural correction interface, so a cheap module can steer each action with fresh vision while the expensive planner keeps running at its own slow rate. It preserves static-task quality while dramatically improving performance on moving targets โ a practical recipe for closing the loop inside the chunk.
Limitations (reviewer): It is a local refinement โ it depends on a reasonable initial plan; when chunks start far from a valid solution, or the target moves near contact, one-step feedback denoising may lack the workspace to recover. Real-world gains are also bounded by occlusion, actuation delay, and observation noise.
5. Links
- arXiv 2609.21022
- Survey: CoRL 2026 ยท Related: Real-Time Execution
โ Back to CoRL 2026 survey ยท Home