ICLR 2026 VLA Robustness - Heungwoo/research GitHub Wiki

RobustVLA — VLA Robustness Against Multi-Modal Perturbations

Venue: ICLR 2026 Affiliation: Beihang · PKU (Psibot Lab) · CUHK · Tsinghua · Zhongguancun Lab · Hefei Comprehensive National Science Center Category: VLA Training — Robustness / safety Trend tag: Multi-modal robustness · adversarial training · UCB bandit

Approach diagram

flowchart LR
  subgraph EVAL[Step 1: Evaluation — 17 perturbations across 4 modalities]
    A[Action perturbations<br/>uniform · Gaussian · bias · flips · spikes]
    O[Observation perturbations<br/>Gaussian · dead pixel · motion blur · color jitter · rotation · shift]
    E[Environment perturbations<br/>external force · irrelevant objects · lighting]
    I[Instruction perturbations<br/>lexical · syntactic · adversarial prompts]
  end
  EVAL --> F1[Finding 1: actions most fragile]
  EVAL --> F2[Finding 2: visual-robust methods do not transfer]
  EVAL --> F3["Finding 3: π₀ beats π₀-FAST beats OpenVLA in robustness"]
  F3 --> RVLA[RobustVLA on π₀ backbone]
  subgraph METHOD[Step 2: RobustVLA training objective]
    RVLA --> OUT[Output robustness L_out:<br/>PGD worst-case δ on action<br/>TRADES-style flow-matching loss]
    RVLA --> IN["Input robustness L_in:<br/>action consistency under ω_i(o_t)<br/>+ PGD on observation"]
    RVLA --> UCB[UCB bandit over Ω<br/>reward = flow-matching loss gap]
  end
  OUT --> Total[L_total = L_π₀ + λ_out L_out + λ_in L_in]
  IN --> Total
  UCB --> IN
Loading

Problem

Prior robust-VLA work (BYOVLA, GEVRM) focuses only on visual perturbations and relies on heavy external VLMs for sensitivity probing or inpainting, costing roughly 10 s/episode of overhead. Yet at deployment a VLA faces uncertainty across four modalities: actions (sensorimotor noise, actuator wear, communication jitter), observations (sensor noise, camera errors), environment (lighting, external forces, distractors), and instructions (synonyms, ambiguity, dialect). The paper's first contribution is a systematic study showing:

  1. Action is the most fragile modality. π₀'s success rate collapses from 96% to 52.4% at 2.5% noise and to ~0% at 5% noise — in stark contrast to robust-RL settings where 10% action noise is routine. Off-distribution accumulates quadratically with horizon in offline imitation, vs. linearly in online RL.
  2. Visual-robust methods do not transfer. BYOVLA gains +7.3% on Gaussian and +22.3% on dead-pixel observation noise but +0.0% on non-visual modalities; average visual gain itself is only +4.0%.
  3. π₀ is the most robust backbone. π₀ beats OpenVLA by 27.9% and π₀-FAST by 5.1%. Since π₀ and π₀-FAST share the VLM, the gap is attributed to the diffusion / flow-matching action head.

Method (detailed)

The framework formalises VLAs as a POMDP G = ⟨Ψ, S, O, O, A, P, R, γ⟩ and defines a uncertainty set Ω = {Ω_ψ ⊆ Ψ, Ω_o ⊆ O, Ω_a ⊆ A, Ω_p ⊆ P} perturbing instructions, observations, actions, and transitions.

4.1 Output robustness — worst-case action noise in flow matching

For π₀'s rectified-flow action head with A^τ_t = τ A¹_t + (1−τ) A⁰_t, the worst-case ℓ_∞-bounded action noise is derived from the flow-matching objective:

δ ∈ argmax_δ E[‖v_θ(Â^τ, o_t) − u(A^τ|A¹) − δ‖²] s.t. ‖δ‖_∞ ≤ ε_action

Computed via PGD (Madry et al. 2017). The robust loss uses TRADES (Zhang et al. 2019) to balance clean and noisy performance:

L_out = max_‖δ‖ E[‖v_θ(o_t, Â^τ(δ), τ) − u^adv_t(δ)‖²]

The authors give three interpretations: (i) flow matching against both clean and adversarially-perturbed action distributions, (ii) label smoothing preventing overconfident matches to specific actions, (iii) an outlier penaliser that quadratically penalises samples the model cannot fit well. The pilot study (Appendix C.5) confirms flow-matching loss correlates with success rate at r = −0.95, p < 0.05.

For autoregressive VLAs (OpenVLA), perturbations are applied to actions before binning and constrained so the result stays within the original or adjacent bins — preserving worst-case proximity to the correct token.

4.2 Input robustness — action consistency under semantic-preserving variations

For each input perturbation ω_i ∈ Ω, the optimal action should not change because the underlying state has not changed. The objective is:

min_θ max_ω_i E[‖v_θ(A^τ, ω_i(o_t)) − u(A^τ|A_t)‖²]

4.3 UCB bandit over perturbation types

With many perturbation arms ω_i, manually weighting each is brittle. The selection is cast as a multi-armed bandit with the UCB rule:

ω*_i = argmax_i [r_n(ω_i) + α · √(log(n) / ω_i(n))]

Reward = increase in flow-matching loss induced by the perturbation (i.e., how harmful it is for the current model), z-score-normalised via EMA. α = 1.0 by default.

Overall objective

min_θ L_RobustVLA = min_θ L_π₀ + λ_in · L_in + λ_out · L_out

with λ_in = λ_out = 1, observation noise η = 8/255, action noise δ = 0.03.

Hyperparameters

Block RobustVLA on π₀ RobustVLA on OpenVLA
Batch size 32 16
Training steps 30,000 30,000
Action expert tuning Full FT –
VLM tuning LoRA LoRA
adv_epsilon (action / obs) 0.03 / 8/255 0.03 / 8/255
pgd_steps 3 3
pgd_alpha 0.01 / 2/255 0.01 / 2/255
UCB exploration coeff 1.0 1.0
UCB window size 100 100
UCB EMA decay 0.9 0.9
UCB min samples 10 10

The action expert (300M parameter Gemma) is fully fine-tuned; the VLM trunk uses LoRA.

Comprehensive Results

LIBERO, 17 perturbations, π₀ backbone (Table 1, success rate %)

Modality Noise type π₀ DR BYOVLA w/o in w/o out w/o UCB RobustVLA
– Clean 96.0 94.8 95.2 96.0 94.9 94.9 95.5
Action Uniform 63.5 61.2 62.0 67.3 66.8 69.5 69.8
Action Gaussian 31.4 30.1 32.1 36.3 31.9 35.9 36.0
Action Bias 23.0 27.6 21.2 44.9 37.0 41.3 42.3
Action Random flips 52.7 48.5 51.6 56.9 54.0 57.6 58.7
Action Sudden spikes 51.7 54.2 51.1 58.5 52.7 62.3 59.7
Obs Gaussian 51.4 64.1 58.7 55.7 94.5 84.3 93.8
Obs Dead pixel 20.8 54.5 43.1 40.8 90.8 78.5 93.8
Obs Motion blur 93.7 88.7 95.2 94.1 95.0 95.2 95.5
Obs Color jitter 61.7 54.0 54.2 53.9 58.7 62.3 69.5
Obs Image rotation 73.3 85.9 77.7 64.1 94.0 70.7 94.4
Obs Image shift 74.6 72.3 70.4 63.2 89.2 62.3 92.7
Env External force 37.1 31.9 37.3 39.0 37.9 39.0 40.8
Env Irrelevant objects 93.1 80.8 93.7 91.2 89.9 91.2 94.2
Env Lighting variation 94.3 89.0 95.0 94.6 94.3 92.4 95.6
Instr Lexical transform 78.7 78.8 79.5 81.5 77.7 77.1 91.3
Instr Syntactic transform 84.7 72.3 85.8 84.0 82.3 83.6 93.9
Instr Adversarial prompts 79.2 56.7 80.1 75.7 72.2 74.6 80.2
Average 17 perturbations 62.6 61.8 64.0 64.8 71.7 69.3 76.6

RobustVLA improves π₀ by +14.0 points absolute and BYOVLA by +12.6 points (97% relative gain on Color Jitter, +73.0 absolute on Dead Pixel, +12.6 on Lexical Transform). Clean performance drops by only −0.5 points (96.0 → 95.5).

OpenVLA backbone (Table 10, 17 perturbations)

Modality Noise OpenVLA BYOVLA RobustVLA
Action Uniform 25.4 24.2 37.6
Action Gaussian 7.4 8.3 10.1
Action Bias 11.8 12.6 24.9
Action Random flips 21.6 20.3 25.4
Action Sudden spikes 22.2 21.5 28.8
Obs Visual Gaussian 0.8 1.5 60.9
Obs Dead pixel 21.6 25.1 68.9
Obs Color jitter 31.0 37.3 38.1
Obs Image rotation 22.3 26.3 26.6
Obs Image shift 42.9 46.6 47.3
Obs Motion blur 59.3 66.1 80.9
Instr Lexical transform 57.7 55.2 58.7
Instr Syntactic 59.1 69.5 76.1
Instr Adversarial prompts 49.3 62.6 64.5
Env Irrelevant objects 72.3 72.7 77.0
Env Lighting variation 64.4 67.4 64.9
Env External force 15.0 14.6 18.2
Average 34.4 37.2 47.6

Absolute gain: +13.2 vs. OpenVLA, +10.4 vs. BYOVLA.

Efficiency (Fig. 4b)

Method Inference time / episode
π₀ ~11 s
RobustVLA ~11 s
BYOVLA ~557 s (50.6×)

BYOVLA requires multiple VLM forward passes plus inpainting per step; RobustVLA pays no extra inference cost.

Mixed perturbations (Fig. 4c)

One random input perturbation + one random output perturbation, same seed across methods:

  • RobustVLA beats π₀ by +14.5% and BYOVLA by +10.4% (p < 0.001, paired t-test).

LIBERO-long (long-horizon, Appendix C.3)

RobustVLA achieves +19.61% average across the four modalities, with the largest gains in observation (+39.93%) and environment (+12%). Long-horizon amplifies the action-modality fragility, validating that horizon length compounds the offline-data action-error problem.

Noise-level sweep (Fig. 11)

  • Action modality: average gain +5.6%, ranges from +1.3% at 0.5% noise to +8.7% at 5% noise.
  • Observation modality: average gain +23%, ranges from +1.9% at 3.9% noise to +74.7% at 35% noise. Gains grow with noise severity.

Real-world Fairino FR5 robot (Fig. 5-6)

Four tasks at four perturbation modes, 10 trials each (omitting external force for safety):

  1. Pick blue bowl on blue plate.
  2. Pick bowl on randomised-colour plate.
  3. Place bread on plate.
  4. Place green cup next to green plate.

Aggregated over the 4 tasks under perturbations with 25 demos (Fig. 5), RobustVLA surpasses the best baseline by +65.6% in success rate — its strength is concentrated in the low-data regime.

Demonstration-scaling study on Task 1 under perturbations (Fig. 6): π₀ rises from 37.5% (25 demos) → 60.0% (50) → 65.0% (100), while RobustVLA holds 92.5% → 92.5% → 95.0%, a +30% gain over π₀ even at 100 demos.

π₀'s real-world robustness plateaus around 65% even with more data; RobustVLA is already highly robust with only 25 demos and its gains do not saturate.

Ablation Studies

  1. w/o input (Ours w/o in): Trains only on output noise. Loses 11.8 points overall (76.6 → 64.8) but still gains on Dead Pixel and other vision-only perturbations through residual benefit of output robustness.
  2. w/o output (Ours w/o out): Loses 4.9 points; retains observation-side gains.
  3. w/o UCB: Drop of 7.3 points — confirms UCB's role in balancing perturbation difficulty rather than overfitting to the easiest ones.
  4. Domain randomisation (DR): Performs essentially at baseline (61.8 vs. 62.6). DR overfits to easy perturbations and degrades on environment/instruction modalities.
  5. PGD ε_action sweep (Table 8): ε ∈ {0.015, 0.03, 0.06}. All beat baseline; ε = 0.03 is the sweet spot (avg 53.3). ε = 0.06 starts to compromise clean accuracy (95.5 → 93.9).
  6. UCB hyperparameter sweep (Table 9): exp_coef ∈ {1.0, 1.5}, ema ∈ {0.6, 0.9}, window ∈ {100, 200}. Average swings by ≤ 1.4 points — robust default.
  7. Adversarial flow-matching loss as proxy (Appendix C.5): Pearson r = −0.953, p < 0.05 between flow-matching loss and task success under action perturbations, justifying using the loss as a differentiable surrogate for success rate.

Limitations (stated or implied by authors)

The paper does not have a dedicated limitations section but identifies and discusses several practical limits:

  • GEVRM not reproduced for fair comparison due to lack of public code / implementation details.
  • External forces in real-world experiments were omitted because they could physically damage the robot — so real-world environment-perturbation results are only partial.
  • PGD steps trade off compute for robustness. More PGD steps yield stronger robustness but increase GPU memory and training wall-clock; the paper uses 3 steps as a compromise.
  • Long-horizon tasks reveal that action modality remains the hardest — gain in actions is only +8% even on LIBERO-long while observation gains are +39.93%. The diffusion-head + offline-data combination fundamentally limits action robustness.
  • Generalisation to unseen perturbations is shown for external force (not in the training set) but not exhaustively benchmarked.

Significance & Positioning

This paper reframes VLA robustness as a four-modality problem and identifies actions — not vision — as the most fragile modality. Versus prior work:

  • BYOVLA: purely visual, slow (50.6× inference overhead), only addresses observation perturbations. RobustVLA is faster, broader, and gets +10.4% over it on the same backbone.
  • GEVRM: purely visual, model-based planning, requires VLM-scale planner.
  • Robust offline RL (e.g., Yang et al. 2022; Shen et al. 2020): treats only state perturbations, not actions. RobustVLA extends robust offline-RL techniques to VLAs by combining TRADES, PGD adversarial training, and a UCB scheduler.
  • π₀ / π₀.6 / GR00T-N1: RobustVLA is a fine-tuning recipe orthogonal to these backbones — it can wrap π₀, π₀-FAST, OpenVLA, or any flow/diffusion/autoregressive VLA.
  • SP-VLA / Action-aware Dynamic Pruning: RobustVLA preserves the inference-speed wins of the underlying VLA (50× faster than BYOVLA).
  • Robust Parameter Merging: alternative robustness recipe in weight space; RobustVLA operates in input/output distribution space.

The 17-perturbation benchmark itself is a contribution. The finding that the diffusion head is intrinsically more robust than autoregressive token prediction is an architectural argument for flow/diffusion-based VLAs that complements π₀.6.

Links

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️