ICLR 2026 VLA Robustness - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 Affiliation: Beihang · PKU (Psibot Lab) · CUHK · Tsinghua · Zhongguancun Lab · Hefei Comprehensive National Science Center Category: VLA Training — Robustness / safety Trend tag: Multi-modal robustness · adversarial training · UCB bandit
flowchart LR
subgraph EVAL[Step 1: Evaluation — 17 perturbations across 4 modalities]
A[Action perturbations<br/>uniform · Gaussian · bias · flips · spikes]
O[Observation perturbations<br/>Gaussian · dead pixel · motion blur · color jitter · rotation · shift]
E[Environment perturbations<br/>external force · irrelevant objects · lighting]
I[Instruction perturbations<br/>lexical · syntactic · adversarial prompts]
end
EVAL --> F1[Finding 1: actions most fragile]
EVAL --> F2[Finding 2: visual-robust methods do not transfer]
EVAL --> F3["Finding 3: π₀ beats π₀-FAST beats OpenVLA in robustness"]
F3 --> RVLA[RobustVLA on π₀ backbone]
subgraph METHOD[Step 2: RobustVLA training objective]
RVLA --> OUT[Output robustness L_out:<br/>PGD worst-case δ on action<br/>TRADES-style flow-matching loss]
RVLA --> IN["Input robustness L_in:<br/>action consistency under ω_i(o_t)<br/>+ PGD on observation"]
RVLA --> UCB[UCB bandit over Ω<br/>reward = flow-matching loss gap]
end
OUT --> Total[L_total = L_π₀ + λ_out L_out + λ_in L_in]
IN --> Total
UCB --> IN
Prior robust-VLA work (BYOVLA, GEVRM) focuses only on visual perturbations and relies on heavy external VLMs for sensitivity probing or inpainting, costing roughly 10 s/episode of overhead. Yet at deployment a VLA faces uncertainty across four modalities: actions (sensorimotor noise, actuator wear, communication jitter), observations (sensor noise, camera errors), environment (lighting, external forces, distractors), and instructions (synonyms, ambiguity, dialect). The paper's first contribution is a systematic study showing:
- Action is the most fragile modality. π₀'s success rate collapses from 96% to 52.4% at 2.5% noise and to ~0% at 5% noise — in stark contrast to robust-RL settings where 10% action noise is routine. Off-distribution accumulates quadratically with horizon in offline imitation, vs. linearly in online RL.
- Visual-robust methods do not transfer. BYOVLA gains +7.3% on Gaussian and +22.3% on dead-pixel observation noise but +0.0% on non-visual modalities; average visual gain itself is only +4.0%.
- π₀ is the most robust backbone. π₀ beats OpenVLA by 27.9% and π₀-FAST by 5.1%. Since π₀ and π₀-FAST share the VLM, the gap is attributed to the diffusion / flow-matching action head.
The framework formalises VLAs as a POMDP G = ⟨Ψ, S, O, O, A, P, R, γ⟩ and defines a uncertainty set Ω = {Ω_ψ ⊆ Ψ, Ω_o ⊆ O, Ω_a ⊆ A, Ω_p ⊆ P} perturbing instructions, observations, actions, and transitions.
For π₀'s rectified-flow action head with A^τ_t = τ A¹_t + (1−τ) A⁰_t, the worst-case ℓ_∞-bounded action noise is derived from the flow-matching objective:
δ ∈ argmax_δ E[‖v_θ(Â^τ, o_t) − u(A^τ|A¹) − δ‖²] s.t. ‖δ‖_∞ ≤ ε_action
Computed via PGD (Madry et al. 2017). The robust loss uses TRADES (Zhang et al. 2019) to balance clean and noisy performance:
L_out = max_‖δ‖ E[‖v_θ(o_t, Â^τ(δ), τ) − u^adv_t(δ)‖²]
The authors give three interpretations: (i) flow matching against both clean and adversarially-perturbed action distributions, (ii) label smoothing preventing overconfident matches to specific actions, (iii) an outlier penaliser that quadratically penalises samples the model cannot fit well. The pilot study (Appendix C.5) confirms flow-matching loss correlates with success rate at r = −0.95, p < 0.05.
For autoregressive VLAs (OpenVLA), perturbations are applied to actions before binning and constrained so the result stays within the original or adjacent bins — preserving worst-case proximity to the correct token.
For each input perturbation ω_i ∈ Ω, the optimal action should not change because the underlying state has not changed. The objective is:
min_θ max_ω_i E[‖v_θ(A^τ, ω_i(o_t)) − u(A^τ|A_t)‖²]
With many perturbation arms ω_i, manually weighting each is brittle. The selection is cast as a multi-armed bandit with the UCB rule:
ω*_i = argmax_i [r_n(ω_i) + α · √(log(n) / ω_i(n))]
Reward = increase in flow-matching loss induced by the perturbation (i.e., how harmful it is for the current model), z-score-normalised via EMA. α = 1.0 by default.
min_θ L_RobustVLA = min_θ L_π₀ + λ_in · L_in + λ_out · L_out
with λ_in = λ_out = 1, observation noise η = 8/255, action noise δ = 0.03.
| Block | RobustVLA on π₀ | RobustVLA on OpenVLA |
|---|---|---|
| Batch size | 32 | 16 |
| Training steps | 30,000 | 30,000 |
| Action expert tuning | Full FT | – |
| VLM tuning | LoRA | LoRA |
adv_epsilon (action / obs) |
0.03 / 8/255 | 0.03 / 8/255 |
pgd_steps |
3 | 3 |
pgd_alpha |
0.01 / 2/255 | 0.01 / 2/255 |
| UCB exploration coeff | 1.0 | 1.0 |
| UCB window size | 100 | 100 |
| UCB EMA decay | 0.9 | 0.9 |
| UCB min samples | 10 | 10 |
The action expert (300M parameter Gemma) is fully fine-tuned; the VLM trunk uses LoRA.
| Modality | Noise type | π₀ | DR | BYOVLA | w/o in | w/o out | w/o UCB | RobustVLA |
|---|---|---|---|---|---|---|---|---|
| – | Clean | 96.0 | 94.8 | 95.2 | 96.0 | 94.9 | 94.9 | 95.5 |
| Action | Uniform | 63.5 | 61.2 | 62.0 | 67.3 | 66.8 | 69.5 | 69.8 |
| Action | Gaussian | 31.4 | 30.1 | 32.1 | 36.3 | 31.9 | 35.9 | 36.0 |
| Action | Bias | 23.0 | 27.6 | 21.2 | 44.9 | 37.0 | 41.3 | 42.3 |
| Action | Random flips | 52.7 | 48.5 | 51.6 | 56.9 | 54.0 | 57.6 | 58.7 |
| Action | Sudden spikes | 51.7 | 54.2 | 51.1 | 58.5 | 52.7 | 62.3 | 59.7 |
| Obs | Gaussian | 51.4 | 64.1 | 58.7 | 55.7 | 94.5 | 84.3 | 93.8 |
| Obs | Dead pixel | 20.8 | 54.5 | 43.1 | 40.8 | 90.8 | 78.5 | 93.8 |
| Obs | Motion blur | 93.7 | 88.7 | 95.2 | 94.1 | 95.0 | 95.2 | 95.5 |
| Obs | Color jitter | 61.7 | 54.0 | 54.2 | 53.9 | 58.7 | 62.3 | 69.5 |
| Obs | Image rotation | 73.3 | 85.9 | 77.7 | 64.1 | 94.0 | 70.7 | 94.4 |
| Obs | Image shift | 74.6 | 72.3 | 70.4 | 63.2 | 89.2 | 62.3 | 92.7 |
| Env | External force | 37.1 | 31.9 | 37.3 | 39.0 | 37.9 | 39.0 | 40.8 |
| Env | Irrelevant objects | 93.1 | 80.8 | 93.7 | 91.2 | 89.9 | 91.2 | 94.2 |
| Env | Lighting variation | 94.3 | 89.0 | 95.0 | 94.6 | 94.3 | 92.4 | 95.6 |
| Instr | Lexical transform | 78.7 | 78.8 | 79.5 | 81.5 | 77.7 | 77.1 | 91.3 |
| Instr | Syntactic transform | 84.7 | 72.3 | 85.8 | 84.0 | 82.3 | 83.6 | 93.9 |
| Instr | Adversarial prompts | 79.2 | 56.7 | 80.1 | 75.7 | 72.2 | 74.6 | 80.2 |
| Average | 17 perturbations | 62.6 | 61.8 | 64.0 | 64.8 | 71.7 | 69.3 | 76.6 |
RobustVLA improves π₀ by +14.0 points absolute and BYOVLA by +12.6 points (97% relative gain on Color Jitter, +73.0 absolute on Dead Pixel, +12.6 on Lexical Transform). Clean performance drops by only −0.5 points (96.0 → 95.5).
| Modality | Noise | OpenVLA | BYOVLA | RobustVLA |
|---|---|---|---|---|
| Action | Uniform | 25.4 | 24.2 | 37.6 |
| Action | Gaussian | 7.4 | 8.3 | 10.1 |
| Action | Bias | 11.8 | 12.6 | 24.9 |
| Action | Random flips | 21.6 | 20.3 | 25.4 |
| Action | Sudden spikes | 22.2 | 21.5 | 28.8 |
| Obs | Visual Gaussian | 0.8 | 1.5 | 60.9 |
| Obs | Dead pixel | 21.6 | 25.1 | 68.9 |
| Obs | Color jitter | 31.0 | 37.3 | 38.1 |
| Obs | Image rotation | 22.3 | 26.3 | 26.6 |
| Obs | Image shift | 42.9 | 46.6 | 47.3 |
| Obs | Motion blur | 59.3 | 66.1 | 80.9 |
| Instr | Lexical transform | 57.7 | 55.2 | 58.7 |
| Instr | Syntactic | 59.1 | 69.5 | 76.1 |
| Instr | Adversarial prompts | 49.3 | 62.6 | 64.5 |
| Env | Irrelevant objects | 72.3 | 72.7 | 77.0 |
| Env | Lighting variation | 64.4 | 67.4 | 64.9 |
| Env | External force | 15.0 | 14.6 | 18.2 |
| Average | 34.4 | 37.2 | 47.6 |
Absolute gain: +13.2 vs. OpenVLA, +10.4 vs. BYOVLA.
| Method | Inference time / episode |
|---|---|
| π₀ | ~11 s |
| RobustVLA | ~11 s |
| BYOVLA | ~557 s (50.6×) |
BYOVLA requires multiple VLM forward passes plus inpainting per step; RobustVLA pays no extra inference cost.
One random input perturbation + one random output perturbation, same seed across methods:
- RobustVLA beats π₀ by +14.5% and BYOVLA by +10.4% (p < 0.001, paired t-test).
RobustVLA achieves +19.61% average across the four modalities, with the largest gains in observation (+39.93%) and environment (+12%). Long-horizon amplifies the action-modality fragility, validating that horizon length compounds the offline-data action-error problem.
- Action modality: average gain +5.6%, ranges from +1.3% at 0.5% noise to +8.7% at 5% noise.
- Observation modality: average gain +23%, ranges from +1.9% at 3.9% noise to +74.7% at 35% noise. Gains grow with noise severity.
Four tasks at four perturbation modes, 10 trials each (omitting external force for safety):
- Pick blue bowl on blue plate.
- Pick bowl on randomised-colour plate.
- Place bread on plate.
- Place green cup next to green plate.
Aggregated over the 4 tasks under perturbations with 25 demos (Fig. 5), RobustVLA surpasses the best baseline by +65.6% in success rate — its strength is concentrated in the low-data regime.
Demonstration-scaling study on Task 1 under perturbations (Fig. 6): π₀ rises from 37.5% (25 demos) → 60.0% (50) → 65.0% (100), while RobustVLA holds 92.5% → 92.5% → 95.0%, a +30% gain over π₀ even at 100 demos.
π₀'s real-world robustness plateaus around 65% even with more data; RobustVLA is already highly robust with only 25 demos and its gains do not saturate.
- w/o input (Ours w/o in): Trains only on output noise. Loses 11.8 points overall (76.6 → 64.8) but still gains on Dead Pixel and other vision-only perturbations through residual benefit of output robustness.
- w/o output (Ours w/o out): Loses 4.9 points; retains observation-side gains.
- w/o UCB: Drop of 7.3 points — confirms UCB's role in balancing perturbation difficulty rather than overfitting to the easiest ones.
- Domain randomisation (DR): Performs essentially at baseline (61.8 vs. 62.6). DR overfits to easy perturbations and degrades on environment/instruction modalities.
- PGD ε_action sweep (Table 8): ε ∈ {0.015, 0.03, 0.06}. All beat baseline; ε = 0.03 is the sweet spot (avg 53.3). ε = 0.06 starts to compromise clean accuracy (95.5 → 93.9).
- UCB hyperparameter sweep (Table 9): exp_coef ∈ {1.0, 1.5}, ema ∈ {0.6, 0.9}, window ∈ {100, 200}. Average swings by ≤ 1.4 points — robust default.
- Adversarial flow-matching loss as proxy (Appendix C.5): Pearson r = −0.953, p < 0.05 between flow-matching loss and task success under action perturbations, justifying using the loss as a differentiable surrogate for success rate.
The paper does not have a dedicated limitations section but identifies and discusses several practical limits:
- GEVRM not reproduced for fair comparison due to lack of public code / implementation details.
- External forces in real-world experiments were omitted because they could physically damage the robot — so real-world environment-perturbation results are only partial.
- PGD steps trade off compute for robustness. More PGD steps yield stronger robustness but increase GPU memory and training wall-clock; the paper uses 3 steps as a compromise.
- Long-horizon tasks reveal that action modality remains the hardest — gain in actions is only +8% even on LIBERO-long while observation gains are +39.93%. The diffusion-head + offline-data combination fundamentally limits action robustness.
- Generalisation to unseen perturbations is shown for external force (not in the training set) but not exhaustively benchmarked.
This paper reframes VLA robustness as a four-modality problem and identifies actions — not vision — as the most fragile modality. Versus prior work:
- BYOVLA: purely visual, slow (50.6× inference overhead), only addresses observation perturbations. RobustVLA is faster, broader, and gets +10.4% over it on the same backbone.
- GEVRM: purely visual, model-based planning, requires VLM-scale planner.
- Robust offline RL (e.g., Yang et al. 2022; Shen et al. 2020): treats only state perturbations, not actions. RobustVLA extends robust offline-RL techniques to VLAs by combining TRADES, PGD adversarial training, and a UCB scheduler.
- π₀ / π₀.6 / GR00T-N1: RobustVLA is a fine-tuning recipe orthogonal to these backbones — it can wrap π₀, π₀-FAST, OpenVLA, or any flow/diffusion/autoregressive VLA.
- SP-VLA / Action-aware Dynamic Pruning: RobustVLA preserves the inference-speed wins of the underlying VLA (50× faster than BYOVLA).
- Robust Parameter Merging: alternative robustness recipe in weight space; RobustVLA operates in input/output distribution space.
The 17-perturbation benchmark itself is a contribution. The finding that the diffusion head is intrinsically more robust than autoregressive token prediction is an architectural argument for flow/diffusion-based VLAs that complements π₀.6.
- OpenReview: https://openreview.net/forum?id=cS6xizdYD5
- Code: https://github.com/gakakulicc/RobustVLA
- Robust Parameter Merging — weight-space robustness
- π0.6
- SP-VLA
- Action-aware Dynamic Pruning
- Align-Then-Steer — adjacent flow-VLA fine-tuning recipe
- Survey: VLA & Manipulation
← Back to ICLR-2026