CVPR 2026 Boost FAN - Heungwoo/research GitHub Wiki
Venue: CVPR 2026 Category: VLA Training Approach Trend tag: Trend 1
Authors: Haochen Niu, Kanyu Zhang (equal), Shuyu Yin, Peilin Liu, Fei Wen (corresponding) — Shanghai Jiao Tong University; Qinghai Guo — Huawei Technologies
flowchart LR
STATE["state s"] --> POL["policy π(a|s)"]
POL --> KL["KL( π || N(μ(s), Σ(s)) )"]
GAUSS["target Gaussian prior<br/>(unimodal, smooth over FAN)"] --> KL
KL --> REG["FAN-guided regularizer"]
REG --> OBJ["+ SFT or RFT objective"]
VLAs are trained with objectives directly inherited from the linguistic setting (e.g., per-token cross-entropy over discretized action bins), but for each state there exists a feasible action neighborhood (FAN) of near-equivalent actions that yield indistinguishable task progress, not a single correct action. Standard VLA training does not exploit this geometry: it pushes an overly confident "spike" onto the single demonstrated bin, hurting generalization and sample efficiency. This affects both supervised fine-tuning (SFT) and reinforced fine-tuning (RFT).
Introduce a FAN-guided regularizer that shapes the policy's output distribution to match the geometry of the feasible action neighborhood. Concretely, it adds the KL divergence between the policy
The prior is instantiated differently per regime:
- SFT: the Gaussian mean is set dynamically to the policy's own argmax and the covariance to the policy's current variance.
-
RFT: a fixed isotropic Gaussian (
$\Sigma = \sigma^2 I$ ) is used for stability.
The method is architecture-agnostic and preserves the discrete, autoregressive action-token format of VLAs (no structural changes). It is demonstrated on OpenVLA (single-action) and OpenVLA-OFT (chunked action sequences), both Llama2-7B backbones.
Evaluated on ManiSkill (sim, pick-and-place with 15 OOD perturbation types), LIBERO (sim), and a real JAKA 7-DoF manipulator (4 tasks, 30 trials each). Gains hold across both SFT and RFT regimes:
| Setting | Comparison | Result |
|---|---|---|
| SFT, ManiSkill | FAN-SFT vs OpenVLA+SFT | +11.7% in-distribution, +5.2% OOD avg |
| RFT, ManiSkill | FAN-PPO vs PPO | +1.5% to +7.9% across models |
| LIBERO-Spatial | FAN-SFT (OpenVLA-OFT) | 98.8% vs 95.2% baseline |
| Real robot | FAN-SFT vs vanilla | task-dependent gains |
Data-scale analysis shows benefits from ~160 up to ~16K samples, and FAN outperforms label smoothing. Note LIBERO baselines are near saturation (>90%), limiting headroom there.
A clean, plug-in regularizer for the standard VLA fine-tuning loop that improves both SFT and RFT without changing the discrete autoregressive action format. Sits adjacent to Knowledge Insulation (which insulates the VLM from action gradients) and to RFT-style methods like VLA-RFT (which use RL gradients explicitly). Most likely to be adopted as a low-cost upgrade to SFT/RFT pipelines.
- arXiv: 2604.01570
← Back to CVPR-2026