ICLR 2026 VER - Heungwoo/research GitHub Wiki

VER — Vision Expert Transformer for Robot Learning via Foundation Distillation and Dynamic Routing

Venue: ICLR 2026 Authors: UC Berkeley + CMU + HKU + PKU + Stony Brook + UNC-Chapel Hill (Yixiao Wang, Mingxiao Huo, Zhixuan Liang, Yushi Du, Lingfeng Sun, Haotian Lin, Jinghuan Shang, Chensheng Peng, Mohit Bansal, Mingyu Ding, Masayoshi Tomizuka) Category: VLA Architecture — Visual encoder · MoE distillation Trend tag: Visual representation · parameter-efficient adaptation · MoE routing

Approach diagram

flowchart LR
  ImgNet[ImageNet-1K] --> BVT[Base Vision Transformer<br/>ViT layers 1..M=9]
  BVT --> VEL[Vision Expert Library<br/>Last N=3 MoE layers<br/>L=6 experts each]
  VEL --> TS[Teacher-Specific Routers R^n_i<br/>top-K=2 per teacher]
  TS --> D1[DINOv2 head]
  TS --> D2[ViT/MAE head]
  TS --> D3[CLIP head]
  D1 --> Loss["L_distill = α·Lcos + (1-α)·LsL1<br/>+ γ·L_mi mutual info"]
  D2 --> Loss
  D3 --> Loss
  Loss --> Frozen[Frozen experts after pretrain]
  Frozen --> RR["Robot Router<br/>under 0.4% params<br/>PER + Curriculum Top-K Annealing"]
  RobotImg[Robot images] --> RR
  RR --> Pol[Policy head<br/>ViLT / Diffusion / Flow-matching]
  Pol --> Act[Action]
Loading

Problem

Single VFMs (DINOv2, CLIP, SAM, ViT) each excel in narrow domains. Naive feature concatenation is heavy and not task-adaptive. Prior multi-VFM distillation (RADIO, Theia) yields static unified representations with three issues: (i) heterogeneous teacher features are misaligned and a unified rep dilutes model-specific capabilities, (ii) policy heads must extract task-relevant info from a fixed fused rep, (iii) full retraining is needed to add robot-domain knowledge. VER's premise: replace the unified rep with a library of specialized experts + a dynamic patch-wise router that selects relevant experts per task and per location.

Detailed Method

Architecture

  • 12-layer ViT backbone, with last N = 3 layers' FFNs replaced by Mixture-of-Experts (MoE).
  • The first M = 9 unaltered layers = Base Vision Transformer (BVT); last 3 = Vision Expert Library (VEL).
  • VEL: each MoE layer has L = 6 expert MLPs.
  • Three model sizes: VER-T (DeiT-Tiny), VER-S (DeiT-Small), VER-B (ViT-Base).
  • Activation: top K = 2 experts per token.

Routing mechanisms

Teacher-Specific Router R^n_i (one per teacher VFM, used during distillation):

  • y = Σ_l R^n_i(x, l) · E^n_l(x), R^n_i(x, l) = m_l · p_l, p = softmax(z), z = s_1 + ε, [s_1; s_2] = MLP(x), ε ~ N(0, SoftPlus(s_2))
  • Noisy gating with top-K hard mask m_l ∈ {0, 1}.

Patchwise Expert Routing (PER) (used downstream): standard MoE routing applied per patch token, < 0.4% additional params.

Curriculum Top-K Annealing (CTA) — the critical training trick:

  • Initialize K_0 = L (all experts active), linearly anneal to K_min over S training steps: K(s) = max(K_min, ⌊L + 1 − (L + 1 − K_min)·(s/S)⌋)
  • Solves "early collapse": Proposition 1 in paper shows that for inactive experts (m_l = 0), the gradient ∂L/∂z_l = -p_l q is independent of expert output → gradients cannot rescue an early-deactivated expert. CTA forces broad early exploration before sparsifying.

Alternative routing modes evaluated:

  • Framewise Teacher Routing (FTR) — one teacher choice per frame.
  • Layerwise Teacher Routing (LTR) — different teacher choices per layer (e.g. DINOv2-like early layers, CLIP-like late layers).
  • PER (default), PER+CTA (best).
  • Gumbel-Softmax with straight-through estimator for discrete teacher selection during training.

Distillation Training (Stage 1)

  • Teachers: DINOv2, ViT (DeiT-style), CLIP — three foundation models distilled into one library.
  • Loss: L_distill = Σ_i α_i [β·L_cos + (1-β)·L_sL1] with α_i = 1/I, β = 0.9.
  • Mutual information loss L_mi = -Σ I(I, E^n) maximizes mutual info between teacher categorical I and routed experts E^n. Equivalent to H(E^n) − H(E^n | I): first term encourages uniform marginal expert use (load balance); second term encourages teacher-specific expert specialization.
  • L_pretrain = L_distill + γ·L_mi, γ = 0.0005.
  • Initialized from Theia weights, trained on ImageNet-1K for 50 epochs on 4× A6000 GPUs.
  • Schedule: 10% linear warmup, 40% constant LR 0.002, 50% Cosine annealing.

Robot Policy Training (Stage 2)

  • Freeze BVT and all VEL experts.
  • Train only the lightweight Robot Router (PER + CTA), <0.4% of total parameters.
  • Routing replaces the static unified-rep approach of Theia/RADIO; experts the policy "listens to" change per patch and per task.

Optional: Train-from-Scratch (TFS) experts

  • Add additional trainable experts to capture robot-domain knowledge missed by VFMs.
  • Best config: 6 distilled-foundation-model (DFM) + 1 TFS, top-K = 2 (Table 5).

Comprehensive Results

11 manipulation tasks across Franka Kitchen + MetaWorld + Adroit (Table 1, success %)

Model LightOn DoorOpen DoorSlide KnobTurn Microwave BinPick ButtonPress DrawerOpen Hammer Pen Relocate Avg
VC-1 1.6 0.2 14.4 1.2 1.8 66.7 56.0 100.0 93.3 68.0 24.0 42.6
MVP 13.6 5.3 17.8 1.8 4.0 73.3 82.7 100.0 97.3 77.7 26.7 48.7
R3M 67.3 31.2 83.1 35.4 35.8 92.0 68.0 100.0 98.7 73.3 58.7 67.6
RADIO 35.2 19.7 69.2 24.4 25.3 82.7 80.0 100.0 100.0 66.7 45.3 61.3
VIP 61.3 25.2 83.0 44.6 31.3 70.7 76.0 98.7 96.0 73.3 29.3 62.8
Theia-B 58.8 34.1 81.2 47.8 24.8 76.0 82.7 100.0 98.7 78.7 46.7 67.1
VER-B (Ours) 67.2 38.0 85.8 55.3 38.2 93.3 94.7 100.0 97.3 80.0 64.0 74.7

VER-B +7.6 pts over Theia-B, +13.4 pts over RADIO, +7.1 over R3M — SOTA across the 11 tasks averaged.

Cross-policy-head generalization (Table 2)

Model LIBERO (ViLT) LIBERO-OOD (ViLT) cross→bin (FM) cube→cup (FM) cylinder→plate (FM) Real-world pour (Diff.)
Theia-T 0.61 0.58 0.65 0.50 0.70 0.45
VER-T 0.70 0.71 0.95 0.75 0.85 0.90

VER beats Theia across ViLT, flow-matching, and diffusion policy heads — the visual encoder gain is policy-head-agnostic. Real-world pour: 0.45 → 0.90.

Vs fine-tuned VLA baseline (Table 10)

Model cross→bin cube→cup cylinder→plate
Theia 0.65 0.50 0.70
GR00T N1.5 (fine-tuned 20k steps, batch 16) 0.75 0.73 0.70
VER (Ours) 0.95 0.75 0.85

VER (only ~5-82M active params + small flow-matching head) outperforms GR00T-N1.5 — a strong VLA — when both train on 500 demos per task. The argument: a strong, task-adaptive vision encoder + small policy is competitive with or better than full VLA fine-tuning at this data scale.

Distillation quality vs Theia (Table 8, lower is better)

Model TP (M) AP (M) Cos↓ DINOv2 Cos↓ ViT Cos↓ CLIP
Theia-T 5.3 5.3 0.641 0.431 0.651
VER-T 7.0 5.3 0.559 0.398 0.592
Theia-S 20.7 20.7 0.554 0.335 0.587
VER-S 27.7 20.8 0.453 0.299 0.517
Theia-B 81.8 81.8 0.444 0.267 0.521
VER-B 110.1 82.2 0.337 0.226 0.455

VER-S already matches Theia-B on distillation loss with ~4× fewer active params. The expert-library design strictly dominates Theia's unified-rep on distillation fidelity.

Inference efficiency

  • VER-T total latency 1.6 ms (router 0.59 ms + experts 0.69 ms) at K=2, L=6.
  • Diffusion policy on RTX 4090: 0.105 s for both VER and Theia — VER adds no inference penalty.

Ablation Studies

Routing mode (Table 3, mean ± SD over 10 seeds)

Task DINOv2 ViT CLIP FTR LTR PER PER+CTA
pen 78.0±4.7 72.8±9.4 80.0±4.6 81.2±3.8 79.2±6.2 78.0±6.3 80.8±5.3
relocate 38.4±5.7 41.6±6.6 41.2±3.8 41.2±6.0 36.4±5.8 47.6±5.1 56.4±6.9

PER+CTA wins with lower variance across seeds (especially relocate). Single-VFM-frozen routing is strictly worse than learned routing.

Top-K sweep (Table 4, VER-Tiny)

K Active params (M) Relocate Pen Avg
1 4.8 42.7 77.3 60.0
2 5.2 52.0 80.0 66.0
3 5.7 57.3 78.7 68.0

Higher K → higher success at higher compute. K=2 chosen as default for compute trade-off.

MoE configuration (Table 11)

K L Success % Total latency (ms) Router Expert
1 3 44.0 ± 11.9 1.45 0.56 0.55
2 4 48.8 ± 5.3 1.49 0.55 0.60
2 6 69.6 ± 4.8 1.62 0.59 0.69
4 12 66.4 ± 4.1 1.94 0.58 1.01

(2, 6) is the sweet spot — pushing to (4, 12) introduces routing instability and degrades performance.

Architecture depth (Table 12)

M N Distill cos loss Success
7 3 0.561 50.4 ± 12.5
9 3 0.551 69.6 ± 4.8
9 5 0.546 48.8 ± 12.2

Counter-intuitive: deeper VEL (N=5) gives best distillation loss but collapses downstream success. Authors interpret: deeper MoE produces high-dim representations that downstream policies can't navigate. Trade-off favors N=3.

CTA exposure schedule

S ∈ {0, 40, 60, 80} epochs evaluated. S=0 (no CTA) leads to seed-dependent collapse. Best at moderate S (depends on benchmark).

Mixing distilled vs trained-from-scratch experts (Table 5)

#DFM #TFS K Relocate Pen Avg
6 0 2 64.0 80.0 72.0
0 2 2 69.3 74.7 72.0
6 1 2 74.7 82.7 78.7

DFM + TFS combination wins — distilled VFM knowledge complements task-specialized experts.

Patch-feature analyses

  • Norm visualizations: without CTA, large outliers in background; with CTA, suppressed → focus on task-critical regions.
  • Mutual information before/after VEL on robot data (Fig. 7, 14): PER+CTA suppresses background MI while preserving task-region MI; lower avg MI ↔ higher success rate (negative regression slope).
  • Comparison vs Theia features (Fig. 8): Theia attends broadly to robot+objects+background; VER post-routing concentrates only on task-relevant objects (e.g. cross + bin, suppressing the robot patches because proprio is fed separately).

Limitations (as stated by authors)

  1. Router training efficiency. PER+CTA needs ~200-400 epochs on Adroit relocate vs typical 100 epochs; "optimizing router training efficiency is left for future work".
  2. Real-world coverage. Only one real task (FANUC LR Mate teapot pour, 20 demos, 120K diffusion training). Broader real-world validation pending.
  3. Distillation tied to specific VFMs. Currently DINOv2 + ViT + CLIP; integrating SAM, RADIO, or larger DINOv3 untested.
  4. TFS-expert scalability. Adding many TFS experts could shift the routing distribution; more rigorous study deferred.
  5. No formal coverage analysis of why top-K annealing succeeds beyond Proposition 1's gradient argument.

Significance & Positioning

Vs Theia / RADIO (multi-VFM distillation peers). Theia distills into one unified representation; RADIO does similar with broader teacher pool. VER's MoE-based library of experts strictly dominates: lower distillation loss at every size, +7.1-13.4 pt avg success on robot tasks. The conceptual upgrade: "distill into many experts, route at inference" rather than "distill into one rep, hope for the best".

Vs single-VFM encoders (R3M, MVP, VC-1, VIP). All these use a single fixed visual encoder. VER's task-adaptive routing beats all of them on the 11-task bench by 7-32 pts. It demonstrates that which VFM matters depends on the task and even the patch, not on a globally-optimal choice.

Vs MoE in general (Switch, Mixtral, sparse-MoE). VER applies MoE to vision encoders for robotics rather than language modeling. Importantly, the MoE here is teacher-specialization-driven (mutual info regularizer), not capacity-scaling-driven.

Vs π₀ / OpenVLA / GR00T (full VLA stack). VER outperforms fine-tuned GR00T-N1.5 on 3 manipulation tasks (Table 10) using a frozen backbone + tiny router + small policy head. Argues that for many manipulation regimes, vision-encoder quality matters more than full VLA scale — a counterpoint to the "scale the VLA" trend. Of course this only holds at the data scale tested (500 demos / task); at much larger data the VLA may dominate.

Vs Spatial-Forcing / spatial-encoder push. Both Spatial Forcing and VER target visual-encoder quality, but Spatial Forcing aligns features to spatial priors, while VER selects among existing VFMs dynamically. Complementary directions.

Vs HiMoE-VLA (hierarchical MoE). Shares the MoE-for-VLA philosophy but at different layers — HiMoE in the policy LLM, VER in the vision encoder.

The deeper bet: parameter-efficient adaptation via routing (<0.4% of params) is enough — most "adaptation" is about selecting the right pretrained knowledge per task, not learning new features. The reduction of high-norm outliers in task-irrelevant regions is a clean visual diagnostic that this happens.

Links

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️