CVPR 2026 CycleVLA - Heungwoo/research GitHub Wiki
CycleVLA โ Proactive Self-Correcting VLA via Subtask Backtracking and Minimum Bayes Risk Decoding
Venue: CVPR 2026 Category: VLA Inference / self-correction Trend tag: Trend 4 Affiliations: Oxford (CS) + Cambridge (Engineering) Authors: Chenyang Ma, Guangyu Yang, Kai Lu, Shitong Xu, Bill Byrne, Niki Trigoni, Andrew Markham Base VLA: OpenVLA backbone with a diffusion-based action expert (action head extended from 7-dim to 9-dim: + stop signal s_tโ{0,1} and progress p_tโ[0,1])
Approach diagram
flowchart LR
OBS["obs + instruction"] --> POL["VLA policy"]
POL --> PROG["progress estimator"]
PROG --> BOUNDARY["subtask boundary detected?"]
BOUNDARY -- yes --> FAIL_PRED["VLM failure predictor"]
FAIL_PRED -- failure likely --> BACK["backtrack to prior subtask"]
FAIL_PRED -- ok --> NEXT["continue"]
POL --> CANDS["sample N action candidates"]
CANDS --> MBR["Minimum Bayes Risk decoding"]
MBR --> ACT["select robust action"]
Problem
VLAs fail mostly at subtask boundaries โ transitions between "approach", "grasp", "lift", etc. โ because that is where the policy must commit to a discrete decision. Once committed, a wrong decision compounds. Standard inference does not detect this and does not recover.
Method
Two complementary mechanisms:
- Progress-aware subtask backtracking โ the VLA itself predicts subtask progress p_t; when progress crosses ฯ_p=0.9 (a critical subtask transition), a VLM (GPT-5.2, third-person + wrist views, current subtask + subtask list) predicts whether the current execution will fail and decides whether to backtrack to the earliest subtask that restores the missing preconditions. State is restored by reverse-executing recorded delta actions. Up to R=3 retries per subtask.
- Minimum Bayes Risk (MBR) decoding โ at the retry step the policy samples N=8 action-chunk hypotheses via stochastic diffusion decoding and selects the one minimizing expected risk, where risk is the mean Lโ distance over predicted end-effector trajectories (a density/medoid variant selects the action in the densest neighborhood). A zero-shot test-time scaling strategy โ more samples โ better selection.
Results
Evaluated on LIBERO (Spatial/Object/Goal/Long, 50 rollouts ร 3 seeds). CycleVLA reaches 95.3% average vs OpenVLA's 76.5%, with the largest gain on Long-horizon (93.6% vs 53.7%); it also surpasses GR00T N1 (93.9% avg). MBR yields the larger gains on under-trained VLAs (~5.3โ9.9% at 200K/350K steps).
Ablation (Table VI): removing MBR drops to 92.5% (โ2.8); removing the stop-signal + last-action oversampling drops to 91.1% (โ4.2). An always-on-MBR upper bound is 96.9%; using the VLM's failure prediction as a hard cutoff (no recovery) collapses to 79.7% โ confirming backtracking and MBR are complementary. Subtask decomposition has only 3.8% relative error (human eval, Table VII).
Significance
CycleVLA is the most explicit inference-time failure recovery mechanism for VLAs published to date. Most prior work treats VLA inference as a single forward pass; CycleVLA explicitly models the episode as a recoverable Markov process with backtracking. Closest sibling: Latent Policy Barrier (which prevents bad actions before they happen rather than recovering from them).
Links
- arXiv: 2601.02295
Related pages
โ Back to CVPR-2026