ICLR 2026 HyperVLA - Heungwoo/research GitHub Wiki

HyperVLA — Efficient Inference via Hypernetworks

Venue: ICLR 2026 Category: VLA Architecture — Efficiency Trend tag: Trend 1 / efficiency

Approach diagram

flowchart LR
  T[Language instruction + initial image o₀<br/>frozen T5 + frozen DINOv2 class token] --> H[Transformer context encoder<br/>HYPERNETWORK 216M<br/>runs ONCE per task]
  H --> W[Generated weights for<br/>compact ViT policy ~0.1M activated]
  W --> P[Compact policy<br/>runs every control step]
  Obs[Observation] --> P
  P --> A[Action]
Loading

Problem

Billion-parameter VLAs deliver strong generalization but demand significant inference compute, limiting deployment to high-end hardware. Distillation and quantization help but trade accuracy.

Method

Use a hypernetwork (HN): a Transformer context encoder reads the task once — conditioned on the frozen T5 language-instruction embedding plus the frozen DINOv2 class-token embedding of the episode's initial image o₀ — and generates the weights of a small task-specific policy. The HN is not a large VLM; it is a high-capacity Transformer with linear output heads (216M HN params + 86M shared params at training time), and it generates only ~0.1M parameters that are actually activated per control step. The base policy is a Vision Transformer (ViT) with a DINOv2 image encoder, a small Transformer policy head, and a linear action head. At deployment only the compact policy runs per control step; the HN sits idle. This is structurally different from distillation — the compact policy is generated per task rather than learned once. Two key design features: an HN normalization technique and a linear (non-diffusion) action head.

Results

Versus OpenVLA, HyperVLA reduces test-time activated parameters by ~90× and achieves a ~120× inference speedup, while matching or exceeding success rates:

  • SIMPLER (zero-shot): Google Robot pick 58±3% (OpenVLA 10%), Google Robot move 73±1% (OpenVLA 72%), WidowX avg 40±5% (OpenVLA 36%).
  • LIBERO (few-shot adaptation): avg 89% vs Octo 75% / OpenVLA 77% (Spatial 95, Object 94, Goal 92, Long 74).

Ablations (Table 4): removing the vision backbone drops avg from 63% to 31%; removing HN normalization degrades OOD tasks to 31%; replacing the linear action head with diffusion falls to 53% avg.

Significance

A third path (alongside distillation and quantization) for deploying large VLAs. Suggests an interesting research direction: if the hypernetwork can condition on more than the task description (specific robot, environment, user preferences), it could produce true per-deployment specialized policies without fine-tuning.

Limitations

Authors flag as future work: real-robot evaluation (results are simulation-only on SIMPLER/LIBERO), scaling up HN model size, and training on larger/more recent robotic datasets.

Links

Related pages

  • AutoQVLA (alternative efficiency approach)
  • π0.6 (target hardware budget)

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️