ICLR 2026 VER - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 Authors: UC Berkeley + CMU + HKU + PKU + Stony Brook + UNC-Chapel Hill (Yixiao Wang, Mingxiao Huo, Zhixuan Liang, Yushi Du, Lingfeng Sun, Haotian Lin, Jinghuan Shang, Chensheng Peng, Mohit Bansal, Mingyu Ding, Masayoshi Tomizuka) Category: VLA Architecture — Visual encoder · MoE distillation Trend tag: Visual representation · parameter-efficient adaptation · MoE routing
flowchart LR
ImgNet[ImageNet-1K] --> BVT[Base Vision Transformer<br/>ViT layers 1..M=9]
BVT --> VEL[Vision Expert Library<br/>Last N=3 MoE layers<br/>L=6 experts each]
VEL --> TS[Teacher-Specific Routers R^n_i<br/>top-K=2 per teacher]
TS --> D1[DINOv2 head]
TS --> D2[ViT/MAE head]
TS --> D3[CLIP head]
D1 --> Loss["L_distill = α·Lcos + (1-α)·LsL1<br/>+ γ·L_mi mutual info"]
D2 --> Loss
D3 --> Loss
Loss --> Frozen[Frozen experts after pretrain]
Frozen --> RR["Robot Router<br/>under 0.4% params<br/>PER + Curriculum Top-K Annealing"]
RobotImg[Robot images] --> RR
RR --> Pol[Policy head<br/>ViLT / Diffusion / Flow-matching]
Pol --> Act[Action]
Single VFMs (DINOv2, CLIP, SAM, ViT) each excel in narrow domains. Naive feature concatenation is heavy and not task-adaptive. Prior multi-VFM distillation (RADIO, Theia) yields static unified representations with three issues: (i) heterogeneous teacher features are misaligned and a unified rep dilutes model-specific capabilities, (ii) policy heads must extract task-relevant info from a fixed fused rep, (iii) full retraining is needed to add robot-domain knowledge. VER's premise: replace the unified rep with a library of specialized experts + a dynamic patch-wise router that selects relevant experts per task and per location.
- 12-layer ViT backbone, with last N = 3 layers' FFNs replaced by Mixture-of-Experts (MoE).
- The first M = 9 unaltered layers = Base Vision Transformer (BVT); last 3 = Vision Expert Library (VEL).
- VEL: each MoE layer has L = 6 expert MLPs.
- Three model sizes: VER-T (DeiT-Tiny), VER-S (DeiT-Small), VER-B (ViT-Base).
- Activation: top K = 2 experts per token.
Teacher-Specific Router R^n_i (one per teacher VFM, used during distillation):
- y = Σ_l R^n_i(x, l) · E^n_l(x), R^n_i(x, l) = m_l · p_l, p = softmax(z), z = s_1 + ε, [s_1; s_2] = MLP(x), ε ~ N(0, SoftPlus(s_2))
- Noisy gating with top-K hard mask m_l ∈ {0, 1}.
Patchwise Expert Routing (PER) (used downstream): standard MoE routing applied per patch token, < 0.4% additional params.
Curriculum Top-K Annealing (CTA) — the critical training trick:
- Initialize K_0 = L (all experts active), linearly anneal to K_min over S training steps: K(s) = max(K_min, ⌊L + 1 − (L + 1 − K_min)·(s/S)⌋)
- Solves "early collapse": Proposition 1 in paper shows that for inactive experts (m_l = 0), the gradient ∂L/∂z_l = -p_l q is independent of expert output → gradients cannot rescue an early-deactivated expert. CTA forces broad early exploration before sparsifying.
Alternative routing modes evaluated:
- Framewise Teacher Routing (FTR) — one teacher choice per frame.
- Layerwise Teacher Routing (LTR) — different teacher choices per layer (e.g. DINOv2-like early layers, CLIP-like late layers).
- PER (default), PER+CTA (best).
- Gumbel-Softmax with straight-through estimator for discrete teacher selection during training.
- Teachers: DINOv2, ViT (DeiT-style), CLIP — three foundation models distilled into one library.
- Loss: L_distill = Σ_i α_i [β·L_cos + (1-β)·L_sL1] with α_i = 1/I, β = 0.9.
- Mutual information loss L_mi = -Σ I(I, E^n) maximizes mutual info between teacher categorical I and routed experts E^n. Equivalent to H(E^n) − H(E^n | I): first term encourages uniform marginal expert use (load balance); second term encourages teacher-specific expert specialization.
- L_pretrain = L_distill + γ·L_mi, γ = 0.0005.
- Initialized from Theia weights, trained on ImageNet-1K for 50 epochs on 4× A6000 GPUs.
- Schedule: 10% linear warmup, 40% constant LR 0.002, 50% Cosine annealing.
- Freeze BVT and all VEL experts.
- Train only the lightweight Robot Router (PER + CTA), <0.4% of total parameters.
- Routing replaces the static unified-rep approach of Theia/RADIO; experts the policy "listens to" change per patch and per task.
- Add additional trainable experts to capture robot-domain knowledge missed by VFMs.
- Best config: 6 distilled-foundation-model (DFM) + 1 TFS, top-K = 2 (Table 5).
| Model | LightOn | DoorOpen | DoorSlide | KnobTurn | Microwave | BinPick | ButtonPress | DrawerOpen | Hammer | Pen | Relocate | Avg |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| VC-1 | 1.6 | 0.2 | 14.4 | 1.2 | 1.8 | 66.7 | 56.0 | 100.0 | 93.3 | 68.0 | 24.0 | 42.6 |
| MVP | 13.6 | 5.3 | 17.8 | 1.8 | 4.0 | 73.3 | 82.7 | 100.0 | 97.3 | 77.7 | 26.7 | 48.7 |
| R3M | 67.3 | 31.2 | 83.1 | 35.4 | 35.8 | 92.0 | 68.0 | 100.0 | 98.7 | 73.3 | 58.7 | 67.6 |
| RADIO | 35.2 | 19.7 | 69.2 | 24.4 | 25.3 | 82.7 | 80.0 | 100.0 | 100.0 | 66.7 | 45.3 | 61.3 |
| VIP | 61.3 | 25.2 | 83.0 | 44.6 | 31.3 | 70.7 | 76.0 | 98.7 | 96.0 | 73.3 | 29.3 | 62.8 |
| Theia-B | 58.8 | 34.1 | 81.2 | 47.8 | 24.8 | 76.0 | 82.7 | 100.0 | 98.7 | 78.7 | 46.7 | 67.1 |
| VER-B (Ours) | 67.2 | 38.0 | 85.8 | 55.3 | 38.2 | 93.3 | 94.7 | 100.0 | 97.3 | 80.0 | 64.0 | 74.7 |
VER-B +7.6 pts over Theia-B, +13.4 pts over RADIO, +7.1 over R3M — SOTA across the 11 tasks averaged.
| Model | LIBERO (ViLT) | LIBERO-OOD (ViLT) | cross→bin (FM) | cube→cup (FM) | cylinder→plate (FM) | Real-world pour (Diff.) |
|---|---|---|---|---|---|---|
| Theia-T | 0.61 | 0.58 | 0.65 | 0.50 | 0.70 | 0.45 |
| VER-T | 0.70 | 0.71 | 0.95 | 0.75 | 0.85 | 0.90 |
VER beats Theia across ViLT, flow-matching, and diffusion policy heads — the visual encoder gain is policy-head-agnostic. Real-world pour: 0.45 → 0.90.
| Model | cross→bin | cube→cup | cylinder→plate |
|---|---|---|---|
| Theia | 0.65 | 0.50 | 0.70 |
| GR00T N1.5 (fine-tuned 20k steps, batch 16) | 0.75 | 0.73 | 0.70 |
| VER (Ours) | 0.95 | 0.75 | 0.85 |
VER (only ~5-82M active params + small flow-matching head) outperforms GR00T-N1.5 — a strong VLA — when both train on 500 demos per task. The argument: a strong, task-adaptive vision encoder + small policy is competitive with or better than full VLA fine-tuning at this data scale.
| Model | TP (M) | AP (M) | Cos↓ DINOv2 | Cos↓ ViT | Cos↓ CLIP |
|---|---|---|---|---|---|
| Theia-T | 5.3 | 5.3 | 0.641 | 0.431 | 0.651 |
| VER-T | 7.0 | 5.3 | 0.559 | 0.398 | 0.592 |
| Theia-S | 20.7 | 20.7 | 0.554 | 0.335 | 0.587 |
| VER-S | 27.7 | 20.8 | 0.453 | 0.299 | 0.517 |
| Theia-B | 81.8 | 81.8 | 0.444 | 0.267 | 0.521 |
| VER-B | 110.1 | 82.2 | 0.337 | 0.226 | 0.455 |
VER-S already matches Theia-B on distillation loss with ~4× fewer active params. The expert-library design strictly dominates Theia's unified-rep on distillation fidelity.
- VER-T total latency 1.6 ms (router 0.59 ms + experts 0.69 ms) at K=2, L=6.
- Diffusion policy on RTX 4090: 0.105 s for both VER and Theia — VER adds no inference penalty.
| Task | DINOv2 | ViT | CLIP | FTR | LTR | PER | PER+CTA |
|---|---|---|---|---|---|---|---|
| pen | 78.0±4.7 | 72.8±9.4 | 80.0±4.6 | 81.2±3.8 | 79.2±6.2 | 78.0±6.3 | 80.8±5.3 |
| relocate | 38.4±5.7 | 41.6±6.6 | 41.2±3.8 | 41.2±6.0 | 36.4±5.8 | 47.6±5.1 | 56.4±6.9 |
PER+CTA wins with lower variance across seeds (especially relocate). Single-VFM-frozen routing is strictly worse than learned routing.
| K | Active params (M) | Relocate | Pen | Avg |
|---|---|---|---|---|
| 1 | 4.8 | 42.7 | 77.3 | 60.0 |
| 2 | 5.2 | 52.0 | 80.0 | 66.0 |
| 3 | 5.7 | 57.3 | 78.7 | 68.0 |
Higher K → higher success at higher compute. K=2 chosen as default for compute trade-off.
| K | L | Success % | Total latency (ms) | Router | Expert |
|---|---|---|---|---|---|
| 1 | 3 | 44.0 ± 11.9 | 1.45 | 0.56 | 0.55 |
| 2 | 4 | 48.8 ± 5.3 | 1.49 | 0.55 | 0.60 |
| 2 | 6 | 69.6 ± 4.8 | 1.62 | 0.59 | 0.69 |
| 4 | 12 | 66.4 ± 4.1 | 1.94 | 0.58 | 1.01 |
(2, 6) is the sweet spot — pushing to (4, 12) introduces routing instability and degrades performance.
| M | N | Distill cos loss | Success |
|---|---|---|---|
| 7 | 3 | 0.561 | 50.4 ± 12.5 |
| 9 | 3 | 0.551 | 69.6 ± 4.8 |
| 9 | 5 | 0.546 | 48.8 ± 12.2 |
Counter-intuitive: deeper VEL (N=5) gives best distillation loss but collapses downstream success. Authors interpret: deeper MoE produces high-dim representations that downstream policies can't navigate. Trade-off favors N=3.
S ∈ {0, 40, 60, 80} epochs evaluated. S=0 (no CTA) leads to seed-dependent collapse. Best at moderate S (depends on benchmark).
| #DFM | #TFS | K | Relocate | Pen | Avg |
|---|---|---|---|---|---|
| 6 | 0 | 2 | 64.0 | 80.0 | 72.0 |
| 0 | 2 | 2 | 69.3 | 74.7 | 72.0 |
| 6 | 1 | 2 | 74.7 | 82.7 | 78.7 |
DFM + TFS combination wins — distilled VFM knowledge complements task-specialized experts.
- Norm visualizations: without CTA, large outliers in background; with CTA, suppressed → focus on task-critical regions.
- Mutual information before/after VEL on robot data (Fig. 7, 14): PER+CTA suppresses background MI while preserving task-region MI; lower avg MI ↔ higher success rate (negative regression slope).
- Comparison vs Theia features (Fig. 8): Theia attends broadly to robot+objects+background; VER post-routing concentrates only on task-relevant objects (e.g. cross + bin, suppressing the robot patches because proprio is fed separately).
- Router training efficiency. PER+CTA needs ~200-400 epochs on Adroit relocate vs typical 100 epochs; "optimizing router training efficiency is left for future work".
- Real-world coverage. Only one real task (FANUC LR Mate teapot pour, 20 demos, 120K diffusion training). Broader real-world validation pending.
- Distillation tied to specific VFMs. Currently DINOv2 + ViT + CLIP; integrating SAM, RADIO, or larger DINOv3 untested.
- TFS-expert scalability. Adding many TFS experts could shift the routing distribution; more rigorous study deferred.
- No formal coverage analysis of why top-K annealing succeeds beyond Proposition 1's gradient argument.
Vs Theia / RADIO (multi-VFM distillation peers). Theia distills into one unified representation; RADIO does similar with broader teacher pool. VER's MoE-based library of experts strictly dominates: lower distillation loss at every size, +7.1-13.4 pt avg success on robot tasks. The conceptual upgrade: "distill into many experts, route at inference" rather than "distill into one rep, hope for the best".
Vs single-VFM encoders (R3M, MVP, VC-1, VIP). All these use a single fixed visual encoder. VER's task-adaptive routing beats all of them on the 11-task bench by 7-32 pts. It demonstrates that which VFM matters depends on the task and even the patch, not on a globally-optimal choice.
Vs MoE in general (Switch, Mixtral, sparse-MoE). VER applies MoE to vision encoders for robotics rather than language modeling. Importantly, the MoE here is teacher-specialization-driven (mutual info regularizer), not capacity-scaling-driven.
Vs π₀ / OpenVLA / GR00T (full VLA stack). VER outperforms fine-tuned GR00T-N1.5 on 3 manipulation tasks (Table 10) using a frozen backbone + tiny router + small policy head. Argues that for many manipulation regimes, vision-encoder quality matters more than full VLA scale — a counterpoint to the "scale the VLA" trend. Of course this only holds at the data scale tested (500 demos / task); at much larger data the VLA may dominate.
Vs Spatial-Forcing / spatial-encoder push. Both Spatial Forcing and VER target visual-encoder quality, but Spatial Forcing aligns features to spatial priors, while VER selects among existing VFMs dynamically. Complementary directions.
Vs HiMoE-VLA (hierarchical MoE). Shares the MoE-for-VLA philosophy but at different layers — HiMoE in the policy LLM, VER in the vision encoder.
The deeper bet: parameter-efficient adaptation via routing (<0.4% of params) is enough — most "adaptation" is about selecting the right pretrained knowledge per task, not learning new features. The reduction of high-norm outliers in task-irrelevant regions is a clean visual diagnostic that this happens.
- OpenReview: https://openreview.net/forum?id=aoorNQFpM6
- Project page: https://yixiaowang7.github.io/ver_page/
- HiMoE-VLA — hierarchical MoE
- Spatial Forcing — spatial visual encoders
- Survey: VLA & Manipulation
← Back to ICLR-2026