ICLR 2026 Robust Param Merging - Heungwoo/research GitHub Wiki

RETAIN — Robust VLA Fine-tuning via Parameter Merging

Venue: ICLR 2026 Category: VLA Training — Adaptation / robustness Trend tag: Continual learning / robust fine-tuning

Approach diagram

flowchart LR
  PT[Pretrained generalist VLA<br/>θ_pre = π0-FAST-DROID / π0-LIBERO] --> FT[Task-FT or Co-FT<br/>~50-100 demos]
  FT --> THFT[Fine-tuned weights θ_ft]
  PT --> Merge["θ̃ = (1-α) · θ_pre + α · θ_ft<br/>α tuned on val OOD scene"]
  THFT --> Merge
  Merge --> Eval[Three evaluations]
  Eval --> ID[Target task ID]
  Eval --> OOD[Target task OOD]
  Eval --> Gen[Generalist tasks]
  Merge -->|continual| Merge2["θ̃_2 = (1-α) · θ̃_1 + α · θ_ft,2<br/>add next skill"]
Loading

Problem

Generalist VLAs (π0, OpenVLA, GR00T, Gemini Robotics) trained on large multi-task corpora generalise impressively out of the box, but practical deployments still fine-tune them on ~50-100 demonstrations for new tasks. In this low-data regime two failure modes appear simultaneously:

  1. Forgetting: the fine-tuned policy degrades on the generalist tasks it could solve before fine-tuning.
  2. Over-fitting: even on the target task, the policy fails on small variations not seen in the fine-tuning dataset (new object instances, lighting, distractors, viewpoints).

Figure 4 in the paper makes this concrete: standard task-FT gradually destroys generalist performance as gradient steps increase, and the gap between in-distribution (ID) and out-of-distribution (OOD) target-task success widens — i.e. the pretrained policy's generalisation ability is not transferring to the new task.

The authors show that even careful learning-rate / step-count tuning (Figure 16 in the appendix) does not resolve this: lower LR retains more generalist knowledge but underfits OOD on the target task; higher LR achieves ID success but kills OOD and generalist.

Detailed Method

Core proposal: linear weight interpolation

Given pretrained θ_pre and fine-tuned θ_ft, RETAIN produces a final policy by linear interpolation:

θ̃ = (1 − α) · θ_pre + α · θ_ft (Eq. 2)

where α ∈ [0, 1] is a tunable merging coefficient. No additional training, no inference-time overhead.

Co-fine-tuning (RETAIN-co-FT)

When the pretraining dataset (or a subset) is available, the authors fine-tune on a mix of D_pre and D_η first, then merge:

  • RETAIN-task-FT: merge θ_pre with task-only fine-tuned weights.
  • RETAIN-co-FT: merge θ_pre with co-fine-tuned weights. Consistently better (Sec. 6.2): co-FT prevents target-side overfitting; merging then explicitly re-injects pretrained knowledge.

Modality-specific merging

VLAs are vision-language-action stacks (vision encoder θ_v, language model backbone θ_l, action expert θ_a). Allow independent coefficients (Eq. 3):

θ̃_v = (1−α_v) θ_pre,v + α_v θ_ft,v θ̃_l = (1−α_l) θ_pre,l + α_l θ_ft,l θ̃_a = (1−α_a) θ_pre,a + α_a θ_ft,a

A 3D grid sweep on mugs-on-plates (Figure 11) reveals that α_l (language model) has the largest gradient on OOD performance — best at α_l = 0.8 — while α_v = α_a = 1 is optimal. Merging only the language-model parameters matches full-merging performance.

Continual sequential adaptation (Eq. 4)

For a sequence of target tasks T_1, ..., T_N, accumulate merges:

θ̃_n = (1 − α) θ̃_{n−1} + α θ_ft,n

The same merging operator is reused at each stage, building a single growing policy.

Hyperparameter selection

  • α swept ∈ {0.25, 0.5, 0.75} on DROID (with one OOD scene held out as validation; the chosen α is then applied unchanged to other test OOD scenes).
  • LIBERO uses a similar val/test split.

Pretrained policies used

  • DROID experiments: π₀-FAST-DROID (autoregressive next-token transformer, FAST tokenizer, trained on all of DROID + Physical Intelligence robot data).
  • LIBERO experiments: π₀ (flow-based action expert) fine-tuned on LIBERO-{object, spatial, goal, 90}.

Comprehensive Results

Real-world DROID tasks (Figure 7)

  • whiteboard: 50 human-teleop demos, single fixed setup, 5 eraser positions × 2 orientations.
  • plates: 100 demos, 5 plate colours × 2 dish racks × 2 orientations; 80 demos contain training-set distractors.

OOD tests vary backgrounds, object instances, camera angles. Each policy is evaluated 10 trials per scene; OOD has val + 2 test scenes.

Reported trends (paper plots, Figure 7):

  • Baselines (Task-FT, Co-FT, LoRA, FreezeFT) reach 70-80 % ID success but only 30-50 % OOD success.
  • RETAIN-task-FT and RETAIN-co-FT reach >60 % on plates OOD and ~80 % on whiteboard OOD — comparable to ID performance, indicating the policy generalised the new skill rather than memorising it.
  • On generalist evaluations (44 distinct real-world DROID tasks distributed across 9 scenes and 17 different language instructions, per Tables 3-5 / App. A.6.7; LIBERO uses 20 random pretraining tasks, 5 from each of object/spatial/goal/90), RETAIN matches the pretrained π0-FAST-DROID baseline; Task-FT and LoRA degrade noticeably.

The paper's headline number: RETAIN finetuned policies achieve ~40 % higher OOD success on average than the best prior fine-tuning method on real DROID.

LIBERO simulation tasks (Figure 8, averaged over 3 tasks)

Tasks: pot-on-stove, mugs-on-plates, items-into-basket. ~45 demos each (filtered from 50 provided), tested on 3 OOD scenes per task with new initial positions, backgrounds, distractors.

  • Most baselines saturate near-perfect on ID due to LIBERO's relative simplicity.
  • OOD shows the same pattern as DROID — RETAIN improves over Co-FT, though by a smaller margin than DROID. The authors attribute this to the weaker generalist base: their LIBERO π0 was trained on only 5.3 k trajectories / 117 scenes, vs π0-FAST-DROID's 76 k trajectories / 564 scenes.

Pretraining-data scaling (Figure 9, plates task)

Three pretrained policies merged-then-evaluated:

  1. π0-FAST-DROID (full DROID + PI internal data) — biggest RETAIN gain on OOD.
  2. DROID-only (76k episodes).
  3. DROID-subset (20k episodes) — smallest gain.

ID is similar across all three. OOD gain scales monotonically with pretraining data: with the most general base, OOD performance is "nearly as good as ID" — i.e. the policy fully transfers its base generalisation to the new task.

Modality-specific merging analysis (Figure 11)

  • α_l (language model) most influential: best at α_l ≈ 0.8 on mugs-on-plates.
  • α_v, α_a = 1 best: vision encoder and action expert benefit from staying at the fine-tuned values.
  • Merging only language-backbone parameters matches full-merge OOD performance across all three LIBERO tasks.

Continual learning (Figure 12)

Sequential plates → whiteboard:

  • RETAIN preserves performance on plates after a second round of fine-tuning on whiteboard, while co-FT in the same sequential setting regresses.
  • On the second task (whiteboard), RETAIN also outperforms co-FT under both ID and OOD, despite no special replay.

Learning-rate ablation (Figure 16)

Four LRs evaluated every 100 gradient steps. Larger LR → faster overfitting and generalist collapse to ~0 %. Smaller LR → slower forgetting but worse OOD performance (i.e. underfitting). No LR setting alone resolves the problem — only weight merging does.

Parameter-trajectory analysis (Figures 17-19)

  • Cosine similarity of consecutive parameter-difference vectors during fine-tuning is far from 1 ⇒ fine-tuning path is highly non-linear.
  • 2D PCA projection of the parameter trajectory shows oscillation (low LR) or curving (high LR); merged checkpoints land in a different region of parameter space, not on the fine-tuning trajectory.
  • Singular-value decomposition of the difference matrix shows many non-zero singular values, confirming non-linearity in many directions.

Ablation Studies

  • α sweep on LIBERO (Figure 10). OOD performance peaks at intermediate α; α → 0 (pretrained) gives ~0 % task success; α → 1 (fine-tuned) gives full overfitting; in between, merging benefits emerge.
  • Modality-specific merging. Confirms language-backbone is the dominant contributor — informs efficient deployment.
  • task-FT vs co-FT before merging. RETAIN-co-FT > RETAIN-task-FT on generalist evals nearly always; on OOD the gap is smaller. The authors attribute different roles: co-FT prevents target overfitting at training time; merging elicits pretrained knowledge in parameter space.
  • Pretraining-data scaling. OOD gain from RETAIN scales with pretraining size (Figure 9 → Figure 15).
  • Sequential vs single-task. Sequential RETAIN approaches single-task oracle; sequential co-FT does not.
  • No-merge baselines. Comparison set: Task-FT, Co-FT, LoRA, FreezeFT (freeze language backbone, only update vision encoder + action expert head), Scratch.

Limitations

The authors explicitly state:

  1. Mechanistic understanding incomplete. Why parameter merging produces this generalisation transfer is "an interesting area for future work"; the paper offers analysis (non-linear path) but no rigorous theoretical explanation.
  2. Hyperparameter dependence. The merging coefficient α requires tuning. The authors note RETAIN is robust to α in real-world experiments (one validation OOD scene generalises to test OOD scenes) but a heuristic for choosing α a priori is left to future work.

Implicit limitations from the paper:

  • Smaller OOD gains on LIBERO than on DROID, attributed to the weaker LIBERO base — RETAIN's effectiveness depends on having a sufficiently generalist starting checkpoint.
  • All real-world experiments are on a Franka with DROID-style setups; no humanoid or mobile-manipulator results.
  • Continual learning shown for 2 sequential tasks; longer sequences are not evaluated.
  • Action-expert merging (α_a < 1) is not helpful — the method's elegance partly comes from leaving the action head alone, but this also means RETAIN does not directly address VLA action-head specialisation.

Significance & Positioning

RETAIN is, in the authors' words, the first work to investigate and analyze parameter merging for robot policies. The paper makes three contributions:

  1. Empirical: ~40 % absolute OOD improvement on real-robot DROID tasks at zero additional training or inference cost. This is the kind of "free lunch" the field has been looking for in the small-demo fine-tuning regime that dominates practical VLA deployments.
  2. Methodological: modality-specific merging analysis identifies the language-model backbone as the load-bearing parameter group — concrete guidance for memory-constrained deployments.
  3. Conceptual: weight merging scales with the generality of the base model (Figure 9). This connects directly to the π0.6 / π0.7 / GR00T Series story: as pretraining data and base-model generality grow, RETAIN gets more effective. It is not in tension with continued scale-up; it is a multiplier on top.

Within the 2026 robust-fine-tuning cluster:

  • Align-Then-Steer adapts in distribution space (steering output statistics) while RETAIN adapts in parameter space. Both target the same problem with orthogonal mechanisms; combining them is unexplored.
  • VLA Robustness is more about identifying failure modes; RETAIN is one prescriptive answer.
  • Knowledge Insulation addresses forgetting via gradient-space techniques; RETAIN sidesteps the forgetting question entirely by never overwriting θ_pre.
  • Compared to PEFT methods like LoRA, RETAIN does not freeze parameters — it allows full fine-tuning then merges, which empirically beats LoRA in both ID and OOD on DROID.
  • For continual learning, RETAIN provides an alternative to EWC / Progress-and-Compress / experience replay, requiring no Fisher matrix or replay buffer — only the predecessor checkpoint.

The paper's clearest implication for the field: the next round of VLA system design should treat the pretrained checkpoint as a first-class deployable artifact, not a starting point to be overwritten. Cheap parameter merging then gives every new task the equivalent of a "soft reset" to the generalist.

Links

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️