CVPR 2026 CycleVLA - Heungwoo/research GitHub Wiki

CycleVLA โ€” Proactive Self-Correcting VLA via Subtask Backtracking and Minimum Bayes Risk Decoding

Venue: CVPR 2026 Category: VLA Inference / self-correction Trend tag: Trend 4 Affiliations: Oxford (CS) + Cambridge (Engineering) Authors: Chenyang Ma, Guangyu Yang, Kai Lu, Shitong Xu, Bill Byrne, Niki Trigoni, Andrew Markham Base VLA: OpenVLA backbone with a diffusion-based action expert (action head extended from 7-dim to 9-dim: + stop signal s_tโˆˆ{0,1} and progress p_tโˆˆ[0,1])

Approach diagram

flowchart LR
  OBS["obs + instruction"] --> POL["VLA policy"]
  POL --> PROG["progress estimator"]
  PROG --> BOUNDARY["subtask boundary detected?"]
  BOUNDARY -- yes --> FAIL_PRED["VLM failure predictor"]
  FAIL_PRED -- failure likely --> BACK["backtrack to prior subtask"]
  FAIL_PRED -- ok --> NEXT["continue"]
  POL --> CANDS["sample N action candidates"]
  CANDS --> MBR["Minimum Bayes Risk decoding"]
  MBR --> ACT["select robust action"]

Problem

VLAs fail mostly at subtask boundaries โ€” transitions between "approach", "grasp", "lift", etc. โ€” because that is where the policy must commit to a discrete decision. Once committed, a wrong decision compounds. Standard inference does not detect this and does not recover.

Method

Two complementary mechanisms:

  • Progress-aware subtask backtracking โ€” the VLA itself predicts subtask progress p_t; when progress crosses ฯ„_p=0.9 (a critical subtask transition), a VLM (GPT-5.2, third-person + wrist views, current subtask + subtask list) predicts whether the current execution will fail and decides whether to backtrack to the earliest subtask that restores the missing preconditions. State is restored by reverse-executing recorded delta actions. Up to R=3 retries per subtask.
  • Minimum Bayes Risk (MBR) decoding โ€” at the retry step the policy samples N=8 action-chunk hypotheses via stochastic diffusion decoding and selects the one minimizing expected risk, where risk is the mean Lโ‚‚ distance over predicted end-effector trajectories (a density/medoid variant selects the action in the densest neighborhood). A zero-shot test-time scaling strategy โ€” more samples โ†’ better selection.

Results

Evaluated on LIBERO (Spatial/Object/Goal/Long, 50 rollouts ร— 3 seeds). CycleVLA reaches 95.3% average vs OpenVLA's 76.5%, with the largest gain on Long-horizon (93.6% vs 53.7%); it also surpasses GR00T N1 (93.9% avg). MBR yields the larger gains on under-trained VLAs (~5.3โ€“9.9% at 200K/350K steps).

Ablation (Table VI): removing MBR drops to 92.5% (โˆ’2.8); removing the stop-signal + last-action oversampling drops to 91.1% (โˆ’4.2). An always-on-MBR upper bound is 96.9%; using the VLM's failure prediction as a hard cutoff (no recovery) collapses to 79.7% โ€” confirming backtracking and MBR are complementary. Subtask decomposition has only 3.8% relative error (human eval, Table VII).

Significance

CycleVLA is the most explicit inference-time failure recovery mechanism for VLAs published to date. Most prior work treats VLA inference as a single forward pass; CycleVLA explicitly models the episode as a recoverable Markov process with backtracking. Closest sibling: Latent Policy Barrier (which prevents bad actions before they happen rather than recovering from them).

Links

Related pages

โ† Back to CVPR-2026