CVPR 2026 Boost FAN - Heungwoo/research GitHub Wiki

Boost-FAN — Boosting VLA Finetuning with Feasible Action Neighborhood Prior

Venue: CVPR 2026 Category: VLA Training Approach Trend tag: Trend 1

Authors: Haochen Niu, Kanyu Zhang (equal), Shuyu Yin, Peilin Liu, Fei Wen (corresponding) — Shanghai Jiao Tong University; Qinghai Guo — Huawei Technologies

Approach diagram

flowchart LR
  STATE["state s"] --> POL["policy π(a|s)"]
  POL --> KL["KL( π || N(μ(s), Σ(s)) )"]
  GAUSS["target Gaussian prior<br/>(unimodal, smooth over FAN)"] --> KL
  KL --> REG["FAN-guided regularizer"]
  REG --> OBJ["+ SFT or RFT objective"]
Loading

Problem

VLAs are trained with objectives directly inherited from the linguistic setting (e.g., per-token cross-entropy over discretized action bins), but for each state there exists a feasible action neighborhood (FAN) of near-equivalent actions that yield indistinguishable task progress, not a single correct action. Standard VLA training does not exploit this geometry: it pushes an overly confident "spike" onto the single demonstrated bin, hurting generalization and sample efficiency. This affects both supervised fine-tuning (SFT) and reinforced fine-tuning (RFT).

Method

Introduce a FAN-guided regularizer that shapes the policy's output distribution to match the geometry of the feasible action neighborhood. Concretely, it adds the KL divergence between the policy $\pi$ and a target Gaussian $\mathcal{N}(\mu(s), \Sigma(s))$ to the training objective, promoting locally smooth, unimodal predictions around the preferred action direction/magnitude (transforming the over-confident "spike" into a smooth neighborhood). It is not a positive/negative classification of proposals.

The prior is instantiated differently per regime:

  • SFT: the Gaussian mean is set dynamically to the policy's own argmax and the covariance to the policy's current variance.
  • RFT: a fixed isotropic Gaussian ($\Sigma = \sigma^2 I$) is used for stability.

The method is architecture-agnostic and preserves the discrete, autoregressive action-token format of VLAs (no structural changes). It is demonstrated on OpenVLA (single-action) and OpenVLA-OFT (chunked action sequences), both Llama2-7B backbones.

Results

Evaluated on ManiSkill (sim, pick-and-place with 15 OOD perturbation types), LIBERO (sim), and a real JAKA 7-DoF manipulator (4 tasks, 30 trials each). Gains hold across both SFT and RFT regimes:

Setting Comparison Result
SFT, ManiSkill FAN-SFT vs OpenVLA+SFT +11.7% in-distribution, +5.2% OOD avg
RFT, ManiSkill FAN-PPO vs PPO +1.5% to +7.9% across models
LIBERO-Spatial FAN-SFT (OpenVLA-OFT) 98.8% vs 95.2% baseline
Real robot FAN-SFT vs vanilla task-dependent gains

Data-scale analysis shows benefits from ~160 up to ~16K samples, and FAN outperforms label smoothing. Note LIBERO baselines are near saturation (>90%), limiting headroom there.

Significance

A clean, plug-in regularizer for the standard VLA fine-tuning loop that improves both SFT and RFT without changing the discrete autoregressive action format. Sits adjacent to Knowledge Insulation (which insulates the VLM from action gradients) and to RFT-style methods like VLA-RFT (which use RL gradients explicitly). Most likely to be adopted as a low-cost upgrade to SFT/RFT pipelines.

Links

Related pages

← Back to CVPR-2026

⚠️ **GitHub.com Fallback** ⚠️