ICLR 2026 Robust Param Merging - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 Category: VLA Training — Adaptation / robustness Trend tag: Continual learning / robust fine-tuning
flowchart LR
PT[Pretrained generalist VLA<br/>θ_pre = π0-FAST-DROID / π0-LIBERO] --> FT[Task-FT or Co-FT<br/>~50-100 demos]
FT --> THFT[Fine-tuned weights θ_ft]
PT --> Merge["θ̃ = (1-α) · θ_pre + α · θ_ft<br/>α tuned on val OOD scene"]
THFT --> Merge
Merge --> Eval[Three evaluations]
Eval --> ID[Target task ID]
Eval --> OOD[Target task OOD]
Eval --> Gen[Generalist tasks]
Merge -->|continual| Merge2["θ̃_2 = (1-α) · θ̃_1 + α · θ_ft,2<br/>add next skill"]
Generalist VLAs (π0, OpenVLA, GR00T, Gemini Robotics) trained on large multi-task corpora generalise impressively out of the box, but practical deployments still fine-tune them on ~50-100 demonstrations for new tasks. In this low-data regime two failure modes appear simultaneously:
- Forgetting: the fine-tuned policy degrades on the generalist tasks it could solve before fine-tuning.
- Over-fitting: even on the target task, the policy fails on small variations not seen in the fine-tuning dataset (new object instances, lighting, distractors, viewpoints).
Figure 4 in the paper makes this concrete: standard task-FT gradually destroys generalist performance as gradient steps increase, and the gap between in-distribution (ID) and out-of-distribution (OOD) target-task success widens — i.e. the pretrained policy's generalisation ability is not transferring to the new task.
The authors show that even careful learning-rate / step-count tuning (Figure 16 in the appendix) does not resolve this: lower LR retains more generalist knowledge but underfits OOD on the target task; higher LR achieves ID success but kills OOD and generalist.
Given pretrained θ_pre and fine-tuned θ_ft, RETAIN produces a final policy by linear interpolation:
θ̃ = (1 − α) · θ_pre + α · θ_ft (Eq. 2)
where α ∈ [0, 1] is a tunable merging coefficient. No additional training, no inference-time overhead.
When the pretraining dataset (or a subset) is available, the authors fine-tune on a mix of D_pre and D_η first, then merge:
- RETAIN-task-FT: merge θ_pre with task-only fine-tuned weights.
- RETAIN-co-FT: merge θ_pre with co-fine-tuned weights. Consistently better (Sec. 6.2): co-FT prevents target-side overfitting; merging then explicitly re-injects pretrained knowledge.
VLAs are vision-language-action stacks (vision encoder θ_v, language model backbone θ_l, action expert θ_a). Allow independent coefficients (Eq. 3):
θ̃_v = (1−α_v) θ_pre,v + α_v θ_ft,v θ̃_l = (1−α_l) θ_pre,l + α_l θ_ft,l θ̃_a = (1−α_a) θ_pre,a + α_a θ_ft,a
A 3D grid sweep on mugs-on-plates (Figure 11) reveals that α_l (language model) has the largest gradient on OOD performance — best at α_l = 0.8 — while α_v = α_a = 1 is optimal. Merging only the language-model parameters matches full-merging performance.
For a sequence of target tasks T_1, ..., T_N, accumulate merges:
θ̃_n = (1 − α) θ̃_{n−1} + α θ_ft,n
The same merging operator is reused at each stage, building a single growing policy.
- α swept ∈ {0.25, 0.5, 0.75} on DROID (with one OOD scene held out as validation; the chosen α is then applied unchanged to other test OOD scenes).
- LIBERO uses a similar val/test split.
- DROID experiments: π₀-FAST-DROID (autoregressive next-token transformer, FAST tokenizer, trained on all of DROID + Physical Intelligence robot data).
- LIBERO experiments: π₀ (flow-based action expert) fine-tuned on LIBERO-{object, spatial, goal, 90}.
- whiteboard: 50 human-teleop demos, single fixed setup, 5 eraser positions × 2 orientations.
- plates: 100 demos, 5 plate colours × 2 dish racks × 2 orientations; 80 demos contain training-set distractors.
OOD tests vary backgrounds, object instances, camera angles. Each policy is evaluated 10 trials per scene; OOD has val + 2 test scenes.
Reported trends (paper plots, Figure 7):
- Baselines (Task-FT, Co-FT, LoRA, FreezeFT) reach 70-80 % ID success but only 30-50 % OOD success.
- RETAIN-task-FT and RETAIN-co-FT reach >60 % on plates OOD and ~80 % on whiteboard OOD — comparable to ID performance, indicating the policy generalised the new skill rather than memorising it.
- On generalist evaluations (44 distinct real-world DROID tasks distributed across 9 scenes and 17 different language instructions, per Tables 3-5 / App. A.6.7; LIBERO uses 20 random pretraining tasks, 5 from each of object/spatial/goal/90), RETAIN matches the pretrained π0-FAST-DROID baseline; Task-FT and LoRA degrade noticeably.
The paper's headline number: RETAIN finetuned policies achieve ~40 % higher OOD success on average than the best prior fine-tuning method on real DROID.
Tasks: pot-on-stove, mugs-on-plates, items-into-basket. ~45 demos each (filtered from 50 provided), tested on 3 OOD scenes per task with new initial positions, backgrounds, distractors.
- Most baselines saturate near-perfect on ID due to LIBERO's relative simplicity.
- OOD shows the same pattern as DROID — RETAIN improves over Co-FT, though by a smaller margin than DROID. The authors attribute this to the weaker generalist base: their LIBERO π0 was trained on only 5.3 k trajectories / 117 scenes, vs π0-FAST-DROID's 76 k trajectories / 564 scenes.
Three pretrained policies merged-then-evaluated:
- π0-FAST-DROID (full DROID + PI internal data) — biggest RETAIN gain on OOD.
- DROID-only (76k episodes).
- DROID-subset (20k episodes) — smallest gain.
ID is similar across all three. OOD gain scales monotonically with pretraining data: with the most general base, OOD performance is "nearly as good as ID" — i.e. the policy fully transfers its base generalisation to the new task.
- α_l (language model) most influential: best at α_l ≈ 0.8 on mugs-on-plates.
- α_v, α_a = 1 best: vision encoder and action expert benefit from staying at the fine-tuned values.
- Merging only language-backbone parameters matches full-merge OOD performance across all three LIBERO tasks.
Sequential plates → whiteboard:
- RETAIN preserves performance on plates after a second round of fine-tuning on whiteboard, while co-FT in the same sequential setting regresses.
- On the second task (whiteboard), RETAIN also outperforms co-FT under both ID and OOD, despite no special replay.
Four LRs evaluated every 100 gradient steps. Larger LR → faster overfitting and generalist collapse to ~0 %. Smaller LR → slower forgetting but worse OOD performance (i.e. underfitting). No LR setting alone resolves the problem — only weight merging does.
- Cosine similarity of consecutive parameter-difference vectors during fine-tuning is far from 1 ⇒ fine-tuning path is highly non-linear.
- 2D PCA projection of the parameter trajectory shows oscillation (low LR) or curving (high LR); merged checkpoints land in a different region of parameter space, not on the fine-tuning trajectory.
- Singular-value decomposition of the difference matrix shows many non-zero singular values, confirming non-linearity in many directions.
- α sweep on LIBERO (Figure 10). OOD performance peaks at intermediate α; α → 0 (pretrained) gives ~0 % task success; α → 1 (fine-tuned) gives full overfitting; in between, merging benefits emerge.
- Modality-specific merging. Confirms language-backbone is the dominant contributor — informs efficient deployment.
- task-FT vs co-FT before merging. RETAIN-co-FT > RETAIN-task-FT on generalist evals nearly always; on OOD the gap is smaller. The authors attribute different roles: co-FT prevents target overfitting at training time; merging elicits pretrained knowledge in parameter space.
- Pretraining-data scaling. OOD gain from RETAIN scales with pretraining size (Figure 9 → Figure 15).
- Sequential vs single-task. Sequential RETAIN approaches single-task oracle; sequential co-FT does not.
- No-merge baselines. Comparison set: Task-FT, Co-FT, LoRA, FreezeFT (freeze language backbone, only update vision encoder + action expert head), Scratch.
The authors explicitly state:
- Mechanistic understanding incomplete. Why parameter merging produces this generalisation transfer is "an interesting area for future work"; the paper offers analysis (non-linear path) but no rigorous theoretical explanation.
- Hyperparameter dependence. The merging coefficient α requires tuning. The authors note RETAIN is robust to α in real-world experiments (one validation OOD scene generalises to test OOD scenes) but a heuristic for choosing α a priori is left to future work.
Implicit limitations from the paper:
- Smaller OOD gains on LIBERO than on DROID, attributed to the weaker LIBERO base — RETAIN's effectiveness depends on having a sufficiently generalist starting checkpoint.
- All real-world experiments are on a Franka with DROID-style setups; no humanoid or mobile-manipulator results.
- Continual learning shown for 2 sequential tasks; longer sequences are not evaluated.
- Action-expert merging (α_a < 1) is not helpful — the method's elegance partly comes from leaving the action head alone, but this also means RETAIN does not directly address VLA action-head specialisation.
RETAIN is, in the authors' words, the first work to investigate and analyze parameter merging for robot policies. The paper makes three contributions:
- Empirical: ~40 % absolute OOD improvement on real-robot DROID tasks at zero additional training or inference cost. This is the kind of "free lunch" the field has been looking for in the small-demo fine-tuning regime that dominates practical VLA deployments.
- Methodological: modality-specific merging analysis identifies the language-model backbone as the load-bearing parameter group — concrete guidance for memory-constrained deployments.
- Conceptual: weight merging scales with the generality of the base model (Figure 9). This connects directly to the π0.6 / π0.7 / GR00T Series story: as pretraining data and base-model generality grow, RETAIN gets more effective. It is not in tension with continued scale-up; it is a multiplier on top.
Within the 2026 robust-fine-tuning cluster:
- Align-Then-Steer adapts in distribution space (steering output statistics) while RETAIN adapts in parameter space. Both target the same problem with orthogonal mechanisms; combining them is unexplored.
- VLA Robustness is more about identifying failure modes; RETAIN is one prescriptive answer.
- Knowledge Insulation addresses forgetting via gradient-space techniques; RETAIN sidesteps the forgetting question entirely by never overwriting θ_pre.
- Compared to PEFT methods like LoRA, RETAIN does not freeze parameters — it allows full fine-tuning then merges, which empirically beats LoRA in both ID and OOD on DROID.
- For continual learning, RETAIN provides an alternative to EWC / Progress-and-Compress / experience replay, requiring no Fisher matrix or replay buffer — only the predecessor checkpoint.
The paper's clearest implication for the field: the next round of VLA system design should treat the pretrained checkpoint as a first-class deployable artifact, not a starting point to be overwritten. Cheap parameter merging then gives every new task the equivalent of a "soft reset" to the generalist.
- OpenReview: https://openreview.net/forum?id=uWJwQ5SZoM
- Project: https://retain.yajatyadav.com
- Align-Then-Steer — distribution-space adaptation
- VLA Robustness
- Knowledge Insulation
- π0.6 · π0.7
- Survey: VLA & Manipulation
← Back to ICLR-2026