ICLR 2026 HyperVLA - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 Category: VLA Architecture — Efficiency Trend tag: Trend 1 / efficiency
flowchart LR
T[Language instruction + initial image o₀<br/>frozen T5 + frozen DINOv2 class token] --> H[Transformer context encoder<br/>HYPERNETWORK 216M<br/>runs ONCE per task]
H --> W[Generated weights for<br/>compact ViT policy ~0.1M activated]
W --> P[Compact policy<br/>runs every control step]
Obs[Observation] --> P
P --> A[Action]
Billion-parameter VLAs deliver strong generalization but demand significant inference compute, limiting deployment to high-end hardware. Distillation and quantization help but trade accuracy.
Use a hypernetwork (HN): a Transformer context encoder reads the task once — conditioned on the frozen T5 language-instruction embedding plus the frozen DINOv2 class-token embedding of the episode's initial image o₀ — and generates the weights of a small task-specific policy. The HN is not a large VLM; it is a high-capacity Transformer with linear output heads (216M HN params + 86M shared params at training time), and it generates only ~0.1M parameters that are actually activated per control step. The base policy is a Vision Transformer (ViT) with a DINOv2 image encoder, a small Transformer policy head, and a linear action head. At deployment only the compact policy runs per control step; the HN sits idle. This is structurally different from distillation — the compact policy is generated per task rather than learned once. Two key design features: an HN normalization technique and a linear (non-diffusion) action head.
Versus OpenVLA, HyperVLA reduces test-time activated parameters by ~90× and achieves a ~120× inference speedup, while matching or exceeding success rates:
- SIMPLER (zero-shot): Google Robot pick 58±3% (OpenVLA 10%), Google Robot move 73±1% (OpenVLA 72%), WidowX avg 40±5% (OpenVLA 36%).
- LIBERO (few-shot adaptation): avg 89% vs Octo 75% / OpenVLA 77% (Spatial 95, Object 94, Goal 92, Long 74).
Ablations (Table 4): removing the vision backbone drops avg from 63% to 31%; removing HN normalization degrades OOD tasks to 31%; replacing the linear action head with diffusion falls to 53% avg.
A third path (alongside distillation and quantization) for deploying large VLAs. Suggests an interesting research direction: if the hypernetwork can condition on more than the task description (specific robot, environment, user preferences), it could produce true per-deployment specialized policies without fine-tuning.
Authors flag as future work: real-robot evaluation (results are simulation-only on SIMPLER/LIBERO), scaling up HN model size, and training on larger/more recent robotic datasets.
- arXiv:2510.04898 (Xiong, Li, Wang, Jackson, Foerster, Whiteson — University of Oxford)
- OpenReview (ICLR 2026)
← Back to ICLR-2026